{"id":"9e22a9ba-859d-4c12-b455-396532b1be1c","arxiv_id":"2607.24709","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Structured LLM prompts generate working hand-tracked AR physics simulations in the browser, demonstrated with a wave-and-lamp activity and positive student attitude data.","lead":"A four-part natural-language prompt can make a large language model generate browser AR physics sims controlled by hand gestures, with no coding. Teachers and students can build and iterate these tools, and a small classroom pilot found strong engagement.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The method claim's key evidence is missing: the paper never reports whether the 29 non-coder students in the pilot actually succeeded in generating a working simulation, how many correction iterations they needed, or whether any archived HTML artifact exists.","rationale":"The reader's weakest_assumption bundled two concerns: weak learning evidence (attitude survey, n=29, no control) and LLM-code stability across models/browsers. I agree with both as flagged, but I locate the load-bearing risk in a more specific, unreported quantity: the success rate and iteration burden of the generation loop in the hands of the actual target users. The pilot was a natural experiment for exactly this question and the paper reports nothing about it — only the attitude outcomes. This is the same family as the reader's \"LLM nondeterminism\" point, hence partial rather than full agreement, but it is sharper: not just \"will the code remain stable\" but \"do we have any evidence the loop works for non-coders at all.\" I do not recommend moving the verdict off CONDITIONAL because the reader's conditions (keep limitations explicit, freeze runnable artifacts) already cover the remedy; the concern sharpens which condition is most important — an archived artifact plus a reported multi-model/multi-user success rate would move this toward ACCEPT. The paper is otherwise transparent: full prompts in the appendix, honest limitations paragraph, and a physics-consistency check described. This is a methods demonstration, and the demonstrated capability (authors can do it) is well supported; only the generalization to the claimed audience is under-evidenced.","tokens_in":6642,"tokens_out":1516,"duration_ms":61604,"concrete_test":"Run the exact wave-and-lamp prompt in 20 fresh sessions on each of Gemini, ChatGPT, and Claude (60 runs total), with no author intervention beyond plain-language correction prompts a non-coder could write; record first-run success rate and median iterations to a sim passing the paper's own technical/physical checks. Separately, archive the generated HTML and open it on 3 browsers/OS combos under ordinary room lighting. If first-run success is ≥80% and corrections stay within plain language, claim (a) holds; if not, the method's audience narrows and the paper should say so.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim has two parts: (a) the four-element prompt lets non-coders generate and refine a working AR sim, and (b) the embodied interface has pedagogical promise. The reader's weak-attitude-survey concern covers (b), and the authors concede it. The softer spot is (a). The paper asserts the prompt \"can often be produced in a common LLM tool\" and describes a short iterate-and-fix loop, but gives no reliability data. Crucially, the pilot section is the one place this could have been measured: 29 students \"received the initial prompt, generated the simulation with it, and performed a technical and physical validation.\" Did all 29 succeed? How many correction prompts did they need? Did any fail and get handed a working file? The text is silent. This matters because the paper itself states \"a first prompt is almost never perfect\" — so the method's viability for the target audience (teachers/students without coding background) depends entirely on the refinement loop being executable by users who cannot read the generated MediaPipe/Three.js code to debug it. Additionally, no artifact is frozen: the prompt is given verbatim, but the generated HTML is not archived, the model is \"the latest Pro model... at the time of writing\" (a moving target), and the appendix admits prompts \"may require minor iteration depending on the model.\" If typical first-run failure rates are high or require code-level debugging, the central claim narrows from \"teachers can do this\" to \"the authors can do this.\"","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The authors extend their prior work on LLM-generated browser physics simulations (slider-based) to camera-and-hand-controlled augmented reality. They present a reusable four-element natural-language prompt structure (tools, display, hand controls, optimization) that, when pasted into a frontier LLM, produces a single-file HTML simulation using MediaPipe Hands (and Three.js for 3D overlays). The flagship example maps the thumb-index pinch gap vertically to amplitude/lamp brightness and horizontally to wavelength/hue. Three further simulations (single charge, two charges, right-hand rule) are given with full prompts in the appendix, and the wave-and-lamp activity was piloted with 29 medical-imaging students who completed an attitude survey (means 4.45-4.59 on 5-point items) and open questions. The authors explicitly frame the survey as measuring perceptions and engagement, not learning.","tokens_in":6969,"tokens_out":1925,"duration_ms":84698,"significance":"If the method is as reliable as claimed, this lowers the barrier to embodied AR physics interactives from a rare programming skill set to a natural-language task, which is genuinely significant for practice-oriented PER venues. The manuscript ships several concrete strengths: the full prompts are reproduced verbatim (main text and appendix); the iteration loop is honestly documented, including the two specific failure modes (distance drift, phase discontinuity) and the exact corrective sentences; and Prompt C's handling of mirroring/handedness for the right-hand rule shows real physical and pedagogical care. The pilot limitations (small n, no comparison group, perception rather than learning measures) are candidly acknowledged rather than oversold. The contribution is methodological and demonstrative, not a learning-gain result, and should be evaluated as such.","major_comments":[{"comment":"'In the classroom' section: the pilot is the one place the reliability of the core method claim could have been measured, and the necessary data are absent. The paper's strongest claim is that the four-element prompt lets teachers or students *without coding background* generate and refine a working AR simulation. The text says the 29 students 'received the initial prompt, generated the simulation with it, and performed a technical and a physical validation,' but never reports: how many students produced a working simulation from the initial prompt, how many correction iterations were typically needed, whether any students failed outright or required instructor/code-level intervention, or which model and interface they used. This is load-bearing because the paper itself states 'a first prompt is almost never perfect,' so the method's viability for the target audience depends on the refin","section":"In the classroom"},{"comment":"No artifact is archived, which undermines reproducibility of a methods paper whose subject is a moving target. The prompts are given verbatim (good), but the generated HTML is not deposited (e.g., in a repository or supplementary file), and the model is identified only as 'the latest Pro model available; at the time of writing, this is Gemini 3.1 Pro.' The appendix concedes prompts 'may require minor iteration depending on the model.' Since the central claim is about what a prompt produces, at minimum the working HTML files for the wave-and-lamp simulation and Prompts A-C should be archived with a persistent identifier, and the exact model version and date used for each reported result recorded. Without this, readers cannot verify that the prompt produces the described behavior even at the time of writing.","section":"The prompt / Appendix"},{"comment":"Survey reporting is internally inconsistent and incomplete. The text states the full survey contained nine statements, 'including three about a variant of the simulation in which the same parameters drive a tone instead of a lamp,' but the tone variant is never described in the paper — the reader cannot tell what activity those items refer to or why they were administered. The paper reports only the light-version items, mixing agreement percentages (93%, 86%, 79-100%) with means (4.52, 4.59, 4.45, 4.48) without giving n per item, distributions, or the full instrument. Either present the complete survey (all nine items with the variant described, or remove mention of it) in a table or appendix, and state the response format for the percentages (e.g., top-two-box). As written, a reader cannot assess selection in which results are reported.","section":"In the classroom"}],"minor_comments":[{"comment":"Figures 1 and 2: given that the paper's subject is visual overlay quality, the figure resolution and annotation matter. Please ensure Fig. 1 clearly shows the on-screen amplitude/wavelength readout mentioned in the caption, and consider labeling the gesture axes in Fig. 1.","section":"Figures"},{"comment":"'The prompt' section: the three-step workaround for ChatGPT/Claude ('add an opening sentence... paste the entire prompt') is slightly confusingly worded — step 1 reads as though the sentence is added and then the prompt pasted after it, which is presumably one combined input. Please clarify.","section":"The prompt"},{"comment":"The claim that pinch-and-spread 'is familiar to most students from touchscreen interactions' would benefit from a citation or softening; it is plausible but asserted.","section":"Introduction"},{"comment":"Reference 8-10 access dates (July 20, 2026) and the course year (2026) are consistent with the submission date, but the phrase 'at the time of writing, this is Gemini 3.1 Pro' will age quickly; consider 'as of July 2026' and note model versioning in the reproducibility statement.","section":"References"},{"comment":"The physical check mentions 'a high frequency goes with a short wavelength,' but the simulation as described has no frequency control — the wave travels at constant speed, so frequency is implicit. A sentence clarifying what students observe about frequency would help.","section":"Iterations and validation"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the venue's methods-and-practice scope and is more candid than most AI-in-education demonstrations. However, the central feasibility claim for non-coders currently rests on an unreported classroom data point the authors very likely possess (success/iteration counts from the pilot). I would regard the revision as straightforward if those records exist; if they do not, the paper remains publishable but needs its claims narrowed accordingly. Two of the four authors are on the cited precursor papers (refs 3-4); the self-citation is appropriate given the lineage, not inflated."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real news here is a reusable four-element prompt (tools, display, hand controls, optimization) that turns a plain-language request into a browser AR sim driven by MediaPipe hand tracking. That is a clean, practical extension of the authors’ earlier slider-simulation work. They ship the full wave-and-lamp prompt, three appendix variants (single charge field, two-charge field, unmirrored right-hand rule), and a short iteration narrative that names the exact failure modes (distance drift, wavelength jumps) and the sentences that fixed them. That is reproducible enough for a methods paper in physics education.\n\nWhat they do well: the technical core is honest and usable. Mirroring rules, phase continuity, relative gap measurement, and the deliberate non-mirroring for the right-hand rule show they actually ran the things. Three validation layers (tracking stability, physical color/wavelength order, gesture naturalness) are the right checklist. The classroom piece is modestly framed: n=29 attitude items, high means on “feeling” amplitude and wavelength, open comments that match, and an explicit admission of no comparison group and no learning measures.\n\nSoft spots, in proportion. The stress-test is right on the method claim: the pilot says students “received the initial prompt, generated the simulation… and performed a technical and physical validation,” but never reports success rate, number of correction turns, or whether anyone needed a working file handed to them. Given their own statement that a first prompt is almost never perfect, that omission matters for the “teachers/students with no coding background” audience. No frozen HTML artifact and a moving “latest Pro” model add ordinary reproducibility friction; the appendix already flags model-dependent iteration. The survey itself is fine as engagement data and is not oversold as learning evidence.\n\nCitations are appropriate (PhET/Physlets lineage, their prior prompt papers, Kontra, Nathan, Scherr). No circular fitting. This is for PER people and instructors who want to try embodied interfaces without a graphics pipeline. I would send it to referees; ask them to require a short success/iteration report from the pilot or an archived runnable example, and to keep learning claims limited. Worth engaging if you care about lowering the barrier to AR demos.","headline":"Useful how-to for prompt-built hand-tracked AR physics demos; the method is concrete and the pilot is thin on whether non-coders actually got working code.","tokens_in":7986,"tokens_out":550,"would_cite":true,"duration_ms":12411,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A reusable four-part natural-language prompt can generate browser-based, hand-controlled AR physics simulations that teachers and students create without coding.","keywords":["physics education","augmented reality","generative AI","embodied learning","hand tracking","browser simulations","gesture control"],"falsifier":"A controlled classroom comparison of conceptual gains using the hand-controlled AR wave simulation versus an otherwise identical slider-based version on the same topic, or systematic hand-tracking failure under ordinary classroom lighting and cameras across multiple generated builds.","tokens_in":7848,"feed_emoji":"✋","tokens_out":834,"duration_ms":32616,"temperature":0.7,"pith_summary":"This paper shows that a structured prompt with four fixed elements—tools, display, hand controls, and optimization—can turn a large language model into a maker of working augmented-reality physics simulations. The output is a single HTML file that runs in an ordinary browser: the camera tracks the learner’s hand, and gestures such as pinch-and-spread set physical parameters drawn over the live room. In a pilot class of 29 students, participants reported that controlling amplitude and wavelength with their fingers helped them “feel” those quantities and stay more engaged than in regular learning. The same prompt skeleton yields other demos, from electric-field arrows filling the space around a fingertip to the right-hand rule drawn on the student’s own hand. If the method holds, embodied AR interfaces become something a teacher can design and refine in plain language rather than a specialized software project.","feed_headline":"Prompt builds hand-controlled AR physics sims in-browser","feed_subtitle":"Four fixed prompt parts let teachers skip coding; students pinch to feel wavelength and amplitude.","key_machinery":"The four-element prompt (tools, display, hand controls, optimization). It forces the model to specify the camera and hand-tracking stack, what is drawn and what it represents, which gesture maps to which quantity, and stability rules (mirroring, phase continuity, palm-normalized gap, hold-last-value on lost tracking) so the result is a runnable embodied simulation.","core_discovery":"A structured natural-language prompt with four reusable elements can generate a camera-driven, hand-controlled AR physics simulation as a single browser HTML file, so teachers or students with no coding background can produce it, refine it in plain language, and use it in class; pilot students reported that the pinch-and-spread gesture helped them feel wavelength and amplitude.","pith_inferences":["Students who rehearse finger motions mentally for exams may be converting the gesture mapping into a portable spatial mnemonic, a transfer path sliders rarely create.","Whether to mirror the camera feed is a content decision (handedness) that must be encoded in the prompt, not a cosmetic default.","Palm-normalized gaps and phase-continuity rules look like reusable optimization clauses for any camera-driven educational simulation, not one-off fixes."],"forward_implications":["Teachers can design and iterate AR physics demos in natural language without programmers.","Students can generate, validate, and refine their own embodied simulations as part of the activity.","The same four-element structure extends across topics such as waves, point-charge fields, and the right-hand rule.","Classroom AR physics interfaces become reachable with a camera-equipped browser and free language-model tools.","Local comparison studies of embodied versus slider interfaces become practical because production cost is low."],"fun_headline_variants":["Structured prompt yields hand-controlled AR physics sim in one HTML file","Four-part prompt lets anyone build browser AR physics tools hands-free","Pinch gestures control AR wave sims made from plain-language prompts","Teachers generate camera-driven AR physics demos without coding","Prompt-to-AR pipeline turns text into embodied physics simulations"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That positive attitude scores from one class of twenty-nine students, with no comparison group and no direct learning measures, plus everyday reliability of the generated tracking code, are enough to establish the pedagogical value of the embodied AR interface.","fun_headline_variants_meta":{"raw":{"variants":["Structured prompt yields hand-controlled AR physics sim in one HTML file","Four-part prompt lets anyone build browser AR physics tools hands-free","Pinch gestures control AR wave sims made from plain-language prompts","Teachers generate camera-driven AR physics demos without coding","Prompt-to-AR pipeline turns text into embodied physics simulations"]},"model":"grok-4.5","effort":"low","cost_usd":0.002567,"raw_usage":{"total_tokens":915,"prompt_tokens":622,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":25668000,"prompt_tokens_details":{"text_tokens":622,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":221,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":622,"tokens_out":72,"duration_ms":4276,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T07:09:08.432603+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled classroom comparison of conceptual gains using the hand-controlled AR wave simulation versus an otherwise identical slider-based version on the same topic, or systematic hand-tracking failure under ordinary classroom lighting and cameras across multiple generated builds.","supporting_citations":[],"review_version":1}