{"id":"67fba6e9-db34-4805-bb63-63bc295b32b4","arxiv_id":"2508.14760","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Future elementary teachers gave AI more explicit, audience-aware, and representation-prescriptive prompts when cast as teachers than when cast as students.","lead":"This study compares how 20 future elementary teachers prompt an AI chat tool when acting as teachers versus students. It finds that in the teacher role they more often name the audience, prescribe the representation, and generate a broader variety of representations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dual-role effect is confounded with topic, content familiarity, audience, and task wording; without a crossed or wording-matched design, the observed prompt differences cannot be attributed to teacher versus student role.","rationale":"The reader's weakest-assumption analysis correctly identifies the core threat: the two role conditions differ in topic, audience, content familiarity, and wording, so the observed prompt differences cannot be cleanly attributed to role. My stress-test confirms this is the most load-bearing concern. The paper is a small, honest exploratory qualitative study, and its reported counts are internally consistent with the data shown, but the central inferential claim—that dual role influences prompting strategies—requires a design that separates role from the other bundled variables. The proposed concrete test—a wording-matched, topic-constant crossover—would directly adjudicate this. Because the reader already assigned a CONDITIONAL verdict and this concern is the same one, no verdict change is needed. I am not raising objections about author conduct or framing; the issue is purely about causal identifiability. Independent support is limited: there is no inter-rater reliability, no raw data, and no formal verification, but those are secondary to the main confound. A future crossed design would strengthen the empirical basis, and until then CONDITIONAL is the appropriate verdict.","tokens_in":7499,"tokens_out":2039,"duration_ms":27403,"concrete_test":"Run a within-subject crossover (or a new study with a comparable population) in which each participant completes both role framings on the same physics topic, with identical representation lists and identical part-2 phrasing, and only the role instruction varies. For example, after learning wave velocity, randomly assign half the participants to 'Imagine you are a teacher explaining wave velocity to elementary students' and half to 'Imagine you are a student explaining wave velocity to your English teacher', then have each participant also complete the other role framing in a second session. If the explicit-audience, prescriptive-prompt, and representation-diversity differences persist when topic and wording are held constant, the role effect is supported; if the differences track topic or wording instead, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that acting as a teacher (versus a student) causes pre-service teachers to be more explicit about the target audience, more prescriptive about representation type, and more diverse in their representations. But the two role conditions differ on at least four dimensions simultaneously: (1) physics topic—wave velocity in the teacher role versus Ideal Gas Law in the student role; (2) content familiarity—post-unit in the teacher role versus pre-unit in the student role; (3) audience—elementary students versus an English teacher; and (4) task wording—Figure 1 explicitly says 'Use two of the following representations in your explanations' and 'Repeat everything you have done in part 1', while Figure 2 says 'Now you are interested in how AI would come up with the best representation...' The last difference is especially damaging: the 'prescriptive versus exploratory' distinction (Section IV) may simply reflect the instruction to choose/use representations versus the instruction to ask AI what the best representation would be. Likewise, audience-explicitness may reflect that 'your elementary students' is a more salient pedagogical audience than 'your English teacher' in a hallway conversation, and representation diversity may reflect the topic or the audience's assumed background rather than role. Thus the data are consistent with the role-effect claim but do not uniquely support it. The paper's own limitations (Section V) mention sample size, linear activity, and AI-order bias, but do not address these confounds. The central claim therefore rests on a causal identification that the current design cannot provide.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory qualitative study of 20 elementary pre-service teachers' AI prompting behavior in two roles: as a teacher (explaining wave velocity to elementary students) and as a student (explaining the Ideal Gas Law to an English teacher). Through emergent coding of the prompts, the authors identify three features that differ across roles: explicit mention of the target audience (17 vs. 11 participants), prescriptive versus exploratory prompting (17 vs. 11 participants), and diversity of representation types (5 vs. 4 categories). The authors interpret these differences as evidence that the teacher role fosters more audience-aware, prescriptive, and varied prompting. The manuscript includes a limitations section but does not address the most serious threat to the central claim: the role conditions are confounded with topic, content familiarity, audience, task wording, and session order.","tokens_in":7838,"tokens_out":4134,"duration_ms":45438,"significance":"If the observed effect of role were causally valid, this study would make a useful contribution to pre-service teacher education, AI literacy, and representational competence. The paper addresses a timely and understudied question, uses an authentic classroom setting, and provides detailed exemplar prompts. The within-subject design (the same 20 participants in both roles) is a strength, as is the transparent presentation of raw counts and examples. However, the design cannot uniquely support the causal claim that the teacher versus student role caused the observed differences. The study is best viewed as hypothesis-generating, and the current framing overstates the certainty of the conclusions.","major_comments":[{"comment":"The teacher and student role conditions differ on at least four dimensions besides the intended role: physics topic (wave velocity vs. Ideal Gas Law), content familiarity (post-unit vs. pre-unit), audience (elementary students vs. an English teacher in a hallway), and task wording. Critically, Fig. 1 instructs participants to first choose two representations and then 'repeat everything' with AI, while Fig. 2 asks 'how AI would come up with the best representation.' The higher prescriptiveness in the teacher role could therefore simply be a response to the task instruction, not to the role. Likewise, audience explicitness and representation diversity could be driven by topic or audience salience. Since the research question asks how the roles 'influence' prompting, the causal attribution is not supported. Please reframe the findings as an exploratory comparison and explicitly discuss thes","section":"Section III, Figs. 1 and 2"},{"comment":"The central numerical claims (17 vs. 11 for audience mention; 17 vs. 11 for prescriptive prompts) are based on emergent coding, but no inter-rater reliability, codebook, or second coder is reported. With a sample of 20 and categories that require judgment (e.g., 'prescriptive' vs. 'exploratory'), a single coder's decisions can drive the main results. Please add reliability statistics (e.g., Cohen's kappa) or at least a detailed audit trail. A paired within-subject analysis (e.g., McNemar test) would also clarify whether the differences are robust to individual variation.","section":"Section IV, second paragraph"},{"comment":"The claim that participants produced 'relatively more diverse representations' as teachers rests on 5 vs. 4 representation categories with highly skewed counts (teacher role: pictures 9, cartoons 2, analogies 3, activity 1, graphs 2; student role: pictures 2, analogies 4, story 1, equations 5). No diversity metric is used, categories are not independent (one participant may prompt multiple representations), and the difference is a single category. This is too thin to support the diversity claim, especially because topic and audience may drive representational suitability. Please temper this claim and, if possible, provide per-participant data and a simple diversity index (e.g., Shannon entropy).","section":"Section IV, third paragraph"}],"minor_comments":[{"comment":"The list of features uses '(i)' twice: '... (i) explicit specification ... (i) predetermination ...'. The second should be '(ii)'.","section":"Section IV, first paragraph"},{"comment":"The caption contains a typo: 'Ideas Gas Law' should be 'Ideal Gas Law'.","section":"Figure 2 caption"},{"comment":"In the sentence 'in only two cases the opposite was observed, i.e., representation specified in student role but not specified in the teacher role', 'representation' should likely be 'target audience', based on the surrounding text.","section":"Section IV, first feature"},{"comment":"The limitations paragraph mentions small sample size, linear activity, and between-subject design, but does not acknowledge the role-topic/audience/task-wording confounds. This is the most important limitation and should be stated explicitly.","section":"Section V, Limitations"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely topic with useful qualitative data, but the central causal claim is undermined by confounds that are not acknowledged. The revisions I request—reframing as exploratory, discussing confounds, and adding reliability—are feasible within the manuscript's scope. I would not reject, but the current version overstates the evidence for role as the causal factor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a modest qualitative study that does one genuinely new thing—comparing pre-service elementary teachers' AI prompts in a teacher persona versus a student persona in a physics representation task—and the reported counts are consistent with the claims. But the two conditions differ on more than role: different physics topic, different content familiarity, different audience, and different task wording. That last one is the killer. In the teacher role, Figure 1 explicitly tells them to 'use two of the following representations'; in the student role, Figure 2 asks them to ask AI 'how it would come up with the best representation.' The prescriptive-versus-exploratory finding looks like it could just be a mirror of that wording. The audience-explicitness finding is also fragile—'your elementary students' is a much more concrete pedagogical audience than 'your English teacher' you happen to meet in a hallway. Content familiarity is confounded too: teacher role came after the Waves unit, student role before the Thermodynamics unit. So the data are compatible with a role effect but do not uniquely support it.\n\nWhat the paper does well: the research question is clearly stated, the coding categories are sensible and the exemplar prompts are helpful, and the authors are honest about the small sample and the linear activity. They do not overclaim hugely—the discussion talks about 'tendencies' and 'relatively.' The main gap is that the limitation section mentions sample size and order bias but never the confounds. That should have been flagged.\n\nThe diversity finding is even softer: 5 representation types in teacher role versus 4 in student role, with small counts and overlapping categories. That is a one-category difference; I would not build anything on it.\n\nIf this is framed as an exploratory descriptive study—'prompts differed across two untangled conditions'—it's fine. If the authors want to claim role causes the difference, the design does not support it. I would send it to review because the topic is timely and the comparison is novel enough to be useful, but the referee should push for a discussion of confounds and a softened causal language.","headline":"New comparison, but role is confounded with topic, task wording, and familiarity; worth reviewing as an exploratory study, not as a causal claim.","tokens_in":8274,"tokens_out":2021,"would_cite":false,"duration_ms":24091,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["01.40.Fk"],"model":"deepseek-v4-flash","headline":"When pre-service elementary teachers prompt an AI tool from the teacher's perspective, they name the intended audience, prescribe the representation type, and reach for more representation formats than when they prompt from the student's pe","keywords":["pre-service teachers","generative AI","AI prompting","representational choices","dual role","prescriptive vs exploratory prompts","physics education","science teacher preparation"],"falsifier":"A two-session experiment holding the topic and audience constant while varying only the role framing, or counterbalancing the topics across sessions, would settle the claim. The decisive variant: assign the teacher role to a topic the participant has not yet studied and the student role to a just-completed topic — the paper's role account predicts teacher-role features (audience-naming, prescriptive prompts, diverse representations) still dominate, whereas a familiarity account predicts the opposite.","tokens_in":7445,"feed_emoji":"🧑🏫","tokens_out":16479,"duration_ms":158598,"temperature":0.7,"pith_summary":"This paper asks whether the dual role that pre-service elementary teachers inhabit — student in their own science courses, teacher in their future classrooms — changes how they prompt a generative AI assistant to produce physics representations. The claim it tries to establish is yes: in a two-session comparison of the same 20 people, teacher-role prompts (explaining wave velocity to elementary students) were more explicit about the target audience, more likely to prescribe which representation the AI should generate, and drew on a wider variety of representation types than student-role prompts (explaining the Ideal Gas Law to an English teacher), which tended to be exploratory and let the AI choose. The authors read the teacher-role pattern as a sign of agency and argue that teacher-preparation courses could deliberately invoke the teacher perspective, with AI's multi-modal output as support, to strengthen content learning and representational competence. The paper matters because prompting AI is becoming routine professional work for teachers, and this is one of the first studies to show the same people prompting differently depending on which role they adopt.","feed_headline":"Teacher role makes pre-service teachers more prescriptive with AI","feed_subtitle":"As teachers they name the audience and pick the representation type; as students they let the AI choose.","key_machinery":"The argument is carried by a dual-role, same-subjects contrast: one group of pre-service teachers, two lab sessions, one AI text assistant, and the same task shell — choose two representations from a fixed menu to explain a physics concept to a named audience — with only the role framing changed. The working analytic device is the emergent coding scheme that sorts each prompt along three dimensions: whether the target audience is named, whether the prompt is 'prescriptive' (the user dictates the representation to the AI) or 'exploratory' (the user asks the AI to choose), and which representation types are requested. The paper's conclusions are the between-role differences on these three code","core_discovery":"Stated on the paper's own terms: pre-service elementary teachers' prompting of AI is role-dependent. Session 1 cast participants as teachers, explaining wave velocity to elementary students after completing the waves unit; session 2 cast them as students, asking about the Ideal Gas Law for an English teacher before the thermodynamics unit. Emergent coding of the 20 participants' prompts revealed three features that shifted between roles. Audience explicitness: 17 of 20 named the audience (e.g., '5th graders') in the teacher role versus 11 in the student role. Prescriptiveness: 17 of 20 teacher-role prompts told the AI which representation to produce, versus 11 student-role prompts; and in no","pith_inferences":["The paper's role contrast is bundled with content familiarity: the teacher session followed the waves unit, while the student session preceded the thermodynamics unit. I infer that prescriptiveness may track 'I know this topic' as much as it tracks the teacher hat; a design that counterbalances topics — or holds familiarity fixed while switching the role label — would separate the two.","A second testable extension is a two-by-two design crossing role (teacher vs. student) with familiarity (just taught vs. not yet taught); scoring the three emergent features mechanically on a larger corpus would show whether the pattern is robust beyond 20 participants.","The audience-explicitness gap (17 vs. 11 of 20) is the cheapest feature to act on, so I would expect a simple intervention — having novices state who the explanation is for before they prompt — to reproduce much of the teacher-role benefit even in the student role.","For AI tool designers, the same gap suggests a lightweight prompt-time question of the user ('Who is this for?') whenever the audience is missing from the prompt."],"forward_implications":["Teacher-preparation courses can use the teacher stance as a lever: asking pre-service teachers to approach their own science learning as future teachers appears to elicit more explicit, audience-aware prompting and more diverse representations.","AI literacy training for teachers should treat prompting as role-dependent and practice the two moves the teacher role naturally produces — naming the audience and naming the output format.","Because AI's strength is multimodality, scaffolding pre-service teachers to generate several representations of one concept with AI as a partner is a concrete route into representational competence.","Exploratory prompting in the student role is a legitimate learning behavior: when content knowledge is thin, pre-service teachers sensibly ask the AI to survey options, fitting AI's role as an open-ended tutor."],"supporting_citations":[{"why":"grounds the study in student-generated questions as a meaningful aspect of science learning, letting the authors read prompts as questions","marker":"[5]"},{"why":"documents that pre-service teachers lack a clear understanding of AI capabilities, motivating an empirical look at how they actually prompt","marker":"[12]"},{"why":"supplies the baseline that pre-service teachers ask short task-completion questions, the contrast that makes the prescriptive/exploratory distinction informative","marker":"[15]"},{"why":"the prior dual-role AI study — students using AI for writing and study plans, teachers for lesson planning — that this paper extends to prompting for representations","marker":"[22]"},{"why":"a task-specific prompting-practices study whose sample size and focus the authors invoke for consistency and comparison","marker":"[23]"},{"why":"supplies the construct of representational competence that makes representation choice the appropriate object of study","marker":"[26]"}],"fun_headline_variants":["Role shift changes how teachers prompt AI","Teacher hat makes AI prompts prescriptive","Preservice teachers prompt AI differently by role","In teacher role, AI prompts get more specific","Casting as teacher drives prescriptive AI prompts"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the differences in prompting are caused by the teacher-versus-student role and not by the other things that differed between the two sessions — physics topic, target audience, content familiarity, and task wording; if one of those is the real cause, the central claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["Role shift changes how teachers prompt AI","Teacher hat makes AI prompts prescriptive","Preservice teachers prompt AI differently by role","In teacher role, AI prompts get more specific","Casting as teacher drives prescriptive AI prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1055,"prompt_tokens":711,"completion_tokens":344,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":455,"completion_tokens_details":{"reasoning_tokens":278}},"tokens_in":455,"tokens_out":344,"duration_ms":4775,"temperature":1.0,"reasoning_tokens":278,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:16:58.897870+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A two-session experiment holding the topic and audience constant while varying only the role framing, or counterbalancing the topics across sessions, would settle the claim. The decisive variant: assign the teacher role to a topic the participant has not yet studied and the student role to a just-completed topic — the paper's role account predicts teacher-role features (audience-naming, prescriptive prompts, diverse representations) still dominate, whereas a familiarity account predicts the opposite.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"grounds the study in student-generated questions as a meaningful aspect of science learning, letting the authors read prompts as questions"},{"cited_title":"Chin and D","cited_arxiv_id":null,"evidence_quote":"documents that pre-service teachers lack a clear understanding of AI capabilities, motivating an empirical look at how they actually prompt"},{"cited_title":"Sirnoorkar, P","cited_arxiv_id":null,"evidence_quote":"supplies the baseline that pre-service teachers ask short task-completion questions, the contrast that makes the prescriptive/exploratory distinction informative"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the prior dual-role AI study — students using AI for writing and study plans, teachers for lesson planning — that this paper extends to prompting for representations"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"a task-specific prompting-practices study whose sample size and focus the authors invoke for consistency and comparison"},{"cited_title":"Shafiq, M","cited_arxiv_id":null,"evidence_quote":"supplies the construct of representational competence that makes representation choice the appropriate object of study"}],"review_version":1}