Pith. sign in

REVIEW 4 major objections 3 minor 6 references

A vibe coding learning design to enhance EFL students' talking to, through, and about AI

T0 review · 4 major / 3 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Differences in how students talk to, through, and about AI explain whether vibe coding succeeds.

desk verdict Genuinely useful pilot for teaching vibe coding in EFL, but the causal claim is confounded by platform/access differences; treat the explanation as a hypothesis. read the letter →

arxiv 2509.08854 v1 pith:RMCDEVLK submitted 2025-09-09 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords vibecodingEFLwritingpromptengineeringmetalanguagingAImentalmodelsauthorshipnegotiationdesignthinkingcasestudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a four-hour pilot workshop in which two English-as-a-foreign-language students used natural-language prompts to build their own writing-assistance apps. It proposes that 'vibe coding' succeeds or fails depending on three intertwined ways students use language with AI: prompting it well, negotiating who is authoring the product, and holding a workable mental model of what AI is. The pilot's contrasting cases—one functional app, one static page—are offered as evidence that these metalanguaging differences explain outcome variation. The larger point is that AI can act as a beneficial languaging machine for language learners, and that teachers should explicitly scaffold all three talking modes rather than only teaching prompt tricks.

What carries the argument

The human-AI metalanguaging framework. It names three dimensions—talking to AI (prompt engineering), talking through AI (negotiating authorship and rhetorical load), and talking about AI (mental models of AI)—and treats them as interrelated. The framework is the lens through which the paper reads screen recordings, think-aloud protocols, worksheets, and AI-generated images; the contrast between the two students is meant to show the framework predicts who ships a coherent app.

What would settle it

Run a matched workshop with a dozen or more students, randomly assigning half to structured-prompt training and half to free-form prompting, on the same platform with equal message quotas; the claim would lose support if app functionality and design coherence do not track the training, or if a conversational prompter succeeds while a structured prompter fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that AI acts as a beneficial languaging machine in vibe coding: it provokes learners' language, reflection, and self-regulation. More specifically, the authors argue that three interrelated metalanguaging dimensions—prompt engineering (talking to AI), authorship negotiation (talking through AI), and mental models of AI (talking about AI)—shape whether a learner's natural-language app-building succeeds. In the reported contrast, the student who wrote long, structured, one-shot prompts built a functional app matching her intended design, while the student who used short conversational prompts and encountered platform limits ended with a static page that did not mat

Load-bearing premise

The argument depends on being able to separate the three kinds of talking in the data and on those differences causing the different outcomes, rather than prior skill, platform choice, or technical luck doing the work.

Editorial extensions

If this is right

  • Vibe coding instruction should teach structured prompt engineering: headings, one-shot examples, and planning prompts before entering them.
  • Teachers should ask students to reflect on what their prompt style implies about human versus AI authorship.
  • Students need explicit vocabulary to describe their mental models of AI; simple pre/post image-generation tasks can reveal gaps.
  • Ambitious design ideas should be paired with pragmatic technical planning so platform limits do not derail the whole project.
  • EFL learners can turn authentic writing problems into functional AI-assisted applications within a few hours when the languaging dimensions are supported.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to run the same backward-designed workshop with a larger class, measure the three dimensions separately, and see whether they predict app functionality beyond prior skill and platform luck.
  • The three-dimension lens should transfer outside EFL: prompt style and perceived authorship may serve as a quick diagnostic for why any vibe-coding project stalls.
  • The students' unchanged minimal image prompts suggest that vibe coding alone may not build mental-model vocabulary; naming AI concepts before the task could produce larger shifts in both prompts and design.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This Innovative Practice article introduces a three-dimension 'human-AI metalanguaging' framework—talking to AI (prompt engineering), talking through AI (negotiating authorship), and talking about AI (mental models)—and applies it to a four-hour EFL vibe-coding workshop. Two Form Three students designed AI-assisted writing apps; one (Student A) produced a functional app on POE's App Creator, while the other (Student G) encountered technical barriers on lovable.dev and submitted a partially functional static page. The authors analyze the contrasting cases through worksheets, think-aloud protocols, screen recordings, and AI-generated images, and argue that differences in how students talked to, through, and about AI explain the outcome variations. They conclude that explicit metalanguaging scaffolding should be part of vibe-coding instruction.

Significance. If the proposed framework is validated, it provides a useful vocabulary for AI-mediated language learning and actionable instructional design for EFL vibe-coding. The paper's strengths include a clearly described workshop design, rich multimodal data collection, transparent think-aloud questions, and concrete examples of prompts and app artifacts. The contrasting-case method is appropriate for a hypothesis-generating pilot. However, the central causal claim—that metalanguaging differences explain success and failure—is not supported by the current evidence: the two cases differ in platform, technical barriers, and access, and the analytic categories derive from the same framework used to design the intervention. The paper is a promising proof-of-concept, but its interpretive claims need to be reframed as hypotheses rather than demonstrated causes.

major comments (4)
  1. [Abstract and §4] The claim that differences in talking to, through, and about AI 'explain vibe coding outcome variations' is too strong. Section 4 goes further: Student A's success 'can be attributed directly to how she talked to AI.' With only two self-selected students, no comparison group, and no experimental manipulation, a causal attribution is not warranted. The case is confounded by platform differences: Student G used lovable.dev and was blocked by a free daily message limit, upload errors, and inability to access Gemini (§3.2.3), while Student A used POE's App Creator with no such barriers. The alternative explanation—that Student G would have succeeded with a different platform or unlimited quota—is not ruled out. Please reframe the conclusion in terms of association and hypothesis generation.
  2. [§3] The content analysis lacks a documented coding protocol and inter-rater reliability. The paper states that prompts were analyzed 'in terms of their number, sequence, frequency, length, and content' and that images were 'analyzed and compared,' but no codebook, unitization rules, or reliability indices are provided. Since the central distinctions—structured vs. conversational prompting, high vs. low metalinguistic awareness, mental models as 'precise tool' vs. 'collaborative partner'—are inferential, the absence of reliability evidence weakens the empirical basis for the conclusions. A systematic coding scheme, or at least an explicit acknowledgment that the analysis is illustrative and interpretive, is needed.
  3. [§2.2 and §3–4] There is a circularity concern: the same three-dimension framework was used to design the learning activities and to define the analytic categories for interpreting the outcomes. For example, the workshop scaffolds 'structured prompt engineering,' and the analysis then finds that Student A used structured prompts and succeeded. This does not invalidate the observation, but it means the framework is partly built into the data. The paper should explicitly acknowledge this design/analysis coupling and consider triangulating with alternative analytic lenses or independent outcome measures to strengthen the interpretation.
  4. [§4] The cognitive-load explanation for authorship perceptions is post hoc. Student A planned 85% authorship but reported 30–40%, and Student G reported 80% authorship despite a collaborative mental model. The paper attributes this to 'high cognitive load,' but no cognitive-load measure was taken; the inference rests on self-reported affect and task behavior. This is a plausible hypothesis, but it should be presented as such, not as a demonstrated mechanism. A direct measure (e.g., a validated cognitive-load scale) or additional evidence from think-aloud transcripts would be needed to support the claim.
minor comments (3)
  1. [§3.2.3] The platform name is spelled inconsistently: 'loveable.dev' appears once, while 'lovable.dev' appears later. Please standardize.
  2. [§3.2.4 and §3.1.4] The pre/post draw-a-picture prompts for Student A are identical, and for Student G only one adds 'Upload the picture.' The paper interprets this as 'misunderstood the task or possessed a limited lexicon,' but it may simply reflect a literal response to an under-specified elicitation. Consider softening this interpretation or probing the students' intentions directly.
  3. [General] The paper lacks a dedicated limitations section. Given the small sample, self-selection of high-performing students, and the teacher–student relationship (the first author taught both students), a brief limitations paragraph would help calibrate reader expectations and strengthen the innovative-practice framing.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the three-dimension framework is theory-laden but the outcome contrast is independently assessed; causal overreach is a validity threat, not a circular one.

full rationale

The paper's central derivation—that differences in talking to, through, and about AI explain vibe coding outcomes—is not equivalent to its inputs by construction. The three dimensions are defined independently of the outcome variable: talking to AI is operationalized as prompt length, structure, and sequencing; talking through AI as students' self-reported authorship percentages; talking about AI as image prompts and think-aloud mental-model inferences. The outcome (functional app cohering to intended design) is assessed from the submitted app artifacts and design sketches, not from the framework categories. Thus the success/failure contrast provides independent content, and the framework's use in both workshop design and data interpretation is theory-laden observation, not circular reasoning. No parameter is fitted to the outcome and then relabeled as a prediction; no conclusion is forced by a definitional identity. The self-citations (Woo et al., 2024; Guo & Li, 2024) appear in background and a post-hoc cognitive-load explanation, but they are not load-bearing: removing them would not alter the central claim. The paper's real weakness is causal overreach—Student G's failure is confounded with platform choice and lovable.dev's paywall, so 'attributed directly to how she talked to AI' (Section 4) is not warranted. But that is a confound and validity limitation, not a circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The paper rests on a conceptual framework that is both the intervention and the analysis lens. It assumes the framework's dimensions are real, observable, and causally relevant, and it borrows 'languaging' from Vygotsky and Jones without independent evidence that human-AI interaction has the same properties as human-human languaging.

assumptions (5)
  • domain assumption EFL students can learn to build software by writing natural-language prompts to LLMs during a four-hour workshop.
    Sections 1 and 2.2 assume vibe coding is feasible and meaningful for Form Three EFL students; no prior skill assessment beyond 'strong EFL performers' is reported.
  • ad hoc to paper The three metalanguaging dimensions (talking to, through, and about AI) are distinct, interrelated, and observable in worksheets, videos, and prompts.
    Figure 1 defines the taxonomy, and it is used both as the instructional design and as the analysis lens, so the taxonomy's validity is assumed rather than independently tested.
  • domain assumption Students' prompt engineering strategies reflect their mental models of AI.
    Section 4 infers Student A views AI as a precise tool and Student G as a collaborative partner from prompt length and content alone.
  • domain assumption Reported percentages of authorship are meaningful indicators of 'talking through AI' rather than measurement artifacts.
    The think-aloud protocol asks 'What percentage of the app's ideas are yours/AI's?' and the paper interprets these self-reports as evidence of authorship negotiation despite their small scale and subjectivity.
  • ad hoc to paper Cognitive load can explain the paradox that Student A planned 85% authorship but reported 30-40%, and Student G reported 80% authorship despite a collaborative mental model.
    Section 4 invokes cognitive load, citing Woo et al. (2024), to resolve discrepancies in the data, but cognitive load is not directly measured in this study.
invented entities (1)
  • human-AI metalanguaging framework (talking to, through, and about AI)
    purpose: Analytic lens for designing the workshop and coding student behavior; presented as the paper's central conceptual contribution.
    The framework is introduced by the authors in Figure 1 and then used to interpret the same data it guided; no external validation or independent measure is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A vibe coding learning design to enhance EFL students' talking to, through, and about AI." pith.science (2026). https://pith.science/paper/RMCDEVLK

@misc{pith2026250908854,
  author       = {Pith},
  title        = {Pith review of: A vibe coding learning design to enhance EFL students' talking to, through, and about AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RMCDEVLK}},
  note         = {Machine review of arXiv:2509.08854}
}
read the original abstract

This innovative practice article reports on the piloting of vibe coding (using natural language to create software applications with AI) for English as a Foreign Language (EFL) education. We developed a human-AI meta-languaging framework with three dimensions: talking to AI (prompt engineering), talking through AI (negotiating authorship), and talking about AI (mental models of AI). Using backward design principles, we created a four-hour workshop where two students designed applications addressing authentic EFL writing challenges. We adopted a case study methodology, collecting data from worksheets and video recordings, think-aloud protocols, screen recordings, and AI-generated images. Contrasting cases showed one student successfully vibe coding a functional application cohering to her intended design, while another encountered technical difficulties with major gaps between intended design and actual functionality. Analysis reveals differences in students' prompt engineering approaches, suggesting different AI mental models and tensions in attributing authorship. We argue that AI functions as a beneficial languaging machine, and that differences in how students talk to, through, and about AI explain vibe coding outcome variations. Findings indicate that effective vibe coding instruction requires explicit meta-languaging scaffolding, teaching structured prompt engineering, facilitating critical authorship discussions, and developing vocabulary for articulating AI mental models.

Figures

Figures reproduced from arXiv: 2509.08854 by the authors.

Figure 2
Figure 2. The authentic problem scenario Alt text: Three paragraphs comprising a writing task prompt Students then iteratively created and evaluated their solution. Students chose their own vibe coding software and vibe coded their application on their iPads. We provided students with soft copies of 1) the HKDSE marking guidelines and 2) sample HKDSE English compositions with scores and feedback. We observed students as they … view at source ↗
Figure 3
Figure 3. Student A’s AI-generated representative pictures of AI (left) and AI in education (right) Alt text: Two pictures with the left picture showing three hands touching a digital AI in the center and the right showing people gathered with computers and an AI hologram above them 3.1.2. Inquiring and Analyzing and Developing Ideas Student A conceptualized the design problem as enabling Form 4 students to “learn to practice… view at source ↗
Figure 12
Figure 12. Student G’s AI-generated representative pictures of AI (left) and AI in education (right) Alt text: Two pictures with the left picture showing a humanoid robot and the right showing humanoid robots and students in a classroom 4. Discussion This study demonstrates that in the context of vibe coding, AI can be a beneficial languaging machine (Jones, 2025), provoking EFL students’ language, reflection, and self￾regulat… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

6 extracted references · 6 canonical work pages

  1. [1]

    Significantly, LLMs have mastered human and computer languages and can be equipped with tools to perform tasks autonomously

    Introduction Artificial intelligence (AI) applications for English as a foreign language (EFL) have expanded as teachers and students use large language models (LLMs) with chatbot interfaces like ChatGPT to generate language outputs that support learning (Barrot, 2023). Significantly, LLMs have mastered human and computer languages and can be equipped wit...

  2. [2]

    Materials and Methods 2.1. Context and Participants 4 We developed the innovative practice as a pilot for an annual human–AI creative writing contest delivering AI and English literacy workshops in Hong Kong secondary schools. The first author piloted the 2025-26 contest learning design in a 4-hour English-medium workshop on July 17, 2025, at an all-girls...

  3. [4]

    Simultaneously, the study reveals insights into the complex interplay between how students talk to, through, and about AI

    Discussion This study demonstrates that in the context of vibe coding, AI can be a beneficial languaging machine (Jones, 2025), provoking EFL students’ language, reflection, and self-regulation. Simultaneously, the study reveals insights into the complex interplay between how students talk to, through, and about AI. The differences in how Students A and G...

  4. [5]

    The cases illustrate that successful vibe coding is intertwined with metalanguaging practice: how students talk to, through, and about AI

    Conclusion This innovative practice has demonstrated vibe coding’s potential in EFL contexts, showing how students use natural language to design software for authentic learning scenarios. The cases illustrate that successful vibe coding is intertwined with metalanguaging practice: how students talk to, through, and about AI. Consequently, effective instr...

  5. [6]

    https://so04.tci-thaijo.org/index.php/LEARN/article/view/256743 Guo, K., & Li, D. (2024). Understanding EFL students’ use of self-made AI chatbots as personalized writing assistance tools: A mixed methods study. System, 124, 103362. https://doi.org/10.1016/j.system.2024.103362 Harvard Graduate School of Education. (2022). PZ’s Thinking Routines Toolbox. P...

  6. [2022]

    DT serves as a guiding framework for students to approach vibe coding

    and writing (Wible, 2020). DT serves as a guiding framework for students to approach vibe coding. We specified essential understandings through intended learning outcomes (ILOs). For DT ILOs, we adopted those from the Middle Years Program (MYP) design cycle because it emphasizes creative, critical thinking for real-world problems, and a coherent four-stag...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.