{"id":"363c633c-6486-4f60-a132-e13ae94cc901","arxiv_id":"2507.05616","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An AR headset prototype with human-simulated handwriting input was rated more engaging and comparably easy to use to desktop GeoGebra in a 10-person study.","lead":"This paper builds an augmented reality headset app that turns handwritten math equations into 3D surface plots users can grab and twist in space. A ten-person study found it more engaging than a whiteboard, GeoGebra on a phone, and GeoGebra on a desktop, although the handwriting reading was actually done by a hidden human operator.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The engagement and ease-of-use results were measured against a simulated handwriting recognizer (Sec. 3.3), so the stated contribution—real-time handwritten equation parsing—is not actually tested; the abstract overstates what was evaluated.","rationale":"The reader's weakest_assumption identifies the same concern I would stress: the Wizard-of-Oz implementation in Sec. 3.3 substitutes a human operator for the claimed handwriting-recognition capability. Because every participant-facing interaction with equation input went through this simulated recognizer, the study measures a prototype with perfect recognition, not the actual system contribution. This is a genuine external-validity threat to the abstract's claim that the system features real-time handwritten equation parsing and that this feature drove the reported engagement and ease-of-use results. I do not see this as internal inconsistency; the paper is candid about the WoZ approach and its limitations. A conditional verdict is therefore appropriate: the prototype has merit as a proof of concept, but the headline claim should be reframed and the system should be tested with an actual OCR pipeline before the 'handwritten input' benefit is asserted. Since the reader already reached CONDITIONAL with moderate confidence, my stress-test does not change the verdict.","tokens_in":8255,"tokens_out":2797,"duration_ms":35213,"concrete_test":"Run a controlled follow-up in which the same 10-participant protocol is repeated with a real automatic equation-parsing pipeline (or, if no such pipeline exists, with a WoZ condition that injects realistic recognition errors and latency calibrated from an OCR baseline). If the engagement advantage over GeoGebra Desktop/AR or the 4.6 handwritten-input rating does not survive, the original results are an artifact of the simulated recognizer rather than evidence for the claimed handwriting-input system.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that an HMD AR system with handwritten equation parsing is more engaging, comparably easy to use, and rated most effective at aiding problem-solving. For that claim to hold, the handwriting-input experience that participants evaluated must be representative of the stated capability. It is not: Sec. 3.3 describes a Wizard-of-Oz pipeline in which a human operator watches the headset camera feed and manually enters the equation over a WebSocket, while a loading indicator reading 'OCR Processing...' masks the operator's delay. Participants therefore experienced a recognizer with effectively perfect accuracy, no recognition-error correction loops, and operator-dependent latency. All of the study's input-related ratings—including the 4.6/5 contribution of handwritten equation input to problem-solving and the comparable ease-of-use result vs. GeoGebra Desktop—were obtained under this simulated condition. A real OCR/equation parser will misread handwriting, require correction, and add variable latency; these failure modes are exactly what could erode the measured engagement and ease-of-use advantages. The paper acknowledges this in Sec. 7 ('introduced variability to our evaluation due to inconsistencies in operator performance'), but the abstract and headline findings do not carry that caveat. Thus the central claim is not internally inconsistent; it is externally unverified, and the study's strongest quantitative results are conditional on the WoZ simulation being a faithful stand-in for an automatic recognizer. That is the load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Breaking the Plane, an AR headset application (Meta Quest 3/Unity) intended to let users visualize 3D mathematical surfaces by writing equations by hand, with real-time graph manipulation and a custom 3D function plotter. The handwriting input is implemented as a Wizard-of-Oz pipeline in which a human operator views the headset camera feed and manually enters the equation, with a fake 'OCR Processing...' indicator masking the delay. A within-subjects user study (n=10, multivariate calculus students) compares four conditions: whiteboard, GeoGebra AR, GeoGebra Desktop, and Breaking the Plane, using Likert-scale ratings and ranking questions. The authors report significantly higher engagement for their system over all baselines, ease of use comparable to GeoGebra Desktop, the highest frequency of being ranked most effective at aiding problem-solving, and strong preference for future use.","tokens_in":8456,"tokens_out":5430,"duration_ms":64292,"significance":"If the findings held as stated, the contribution would be a useful demonstration that HMD-based AR with handwriting-style input can rival desktop tools in perceived usability while increasing engagement, addressing the content-authoring friction that prior AR math tools report. The paper has strengths: it compares against three external baselines, uses randomized system-query pairings, uses a within-subjects design, and transparently discloses the Wizard-of-Oz limitation, small sample, and learning effects. However, the empirical basis is thin and partly conditional on a simulated component; the headline claims overstate what was actually tested. As an exploratory extended-abstract study, the work is worth reporting, but the central claims need reframing to match the evidence.","major_comments":[{"comment":"The core input capability, 'handwritten equation parsing,' is not implemented or tested: a human Wizard manually transcribes the equation over a WebSocket, and the user sees a simulated 'OCR Processing...' indicator. All input-related ratings, including the 4.6/5 contribution of handwritten input in §5.4, were collected under a condition with effectively perfect recognition and no error-correction loops. The paper acknowledges this in §7 ('introduced variability to our evaluation due to inconsistencies in operator performance'), but the Abstract and §8 present the system as having 'handwritten equation-parsing input' without this caveat. Since a real OCR parser would add latency, misrecognition, and correction friction—exactly the factors that could erode the measured engagement and ease-of-use advantages—the central claim as stated is not supported. Please either reframe all headline findings as applying to a Wizard-of-Oz-simulated handwriting interface, or supply evidence (e.g., operator transcription accuracy/latency logs and a sensitivity analysis) to justify generalizing to a real recognizer.","section":"§3.3"},{"comment":"The inferential statistics rely on multiple paired t-tests on 5-point Likert items without correction for multiple comparisons. For the engagement comparisons, three pairwise tests are reported; under a simple Bonferroni correction (α = 0.05/3 ≈ 0.0167), the Breaking the Plane vs. GeoGebra AR comparison (p = 0.022) no longer meets the threshold, so the Abstract's 'significantly surpassed other tools' is not robust to standard multiple-comparison control. The ease-of-use section also reports several pairwise tests without familywise error control. Please report corrected p-values or false-discovery-rate adjustments, and include effect sizes or confidence intervals for the pairwise differences.","section":"§5"},{"comment":"Problem-solving effectiveness is measured only by self-report, not by task performance. The study questions were explicitly 'not graded for correctness' (§4.2.1), and the reported evidence is a ranking (6 first, 4 second) plus retrospective Likert ratings. The claim that the system was 'the most effective in aiding problem-solving' should therefore be worded as perceived effectiveness, or supported by objective measures such as the number of correct answers or time-to-solution across conditions.","section":"§4.2.1"}],"minor_comments":[{"comment":"The sentence 'The t-test results provided in Fig. 3 demonstrate significant variances in ease of use' should say 'significant differences' rather than 'variances,' since the analysis concerns means, not variances.","section":"§5.1"},{"comment":"The t-statistics for engagement are all negative; the sign convention should be stated explicitly (e.g., negative values indicate higher mean engagement for 'Our System') so readers can interpret the direction of the effects.","section":"§5.2"},{"comment":"Figure 4, comparing 'our system with GeoGebra Desktop and GeoGebra AR in aiding problem-solving,' is mentioned only briefly in the text; the figure's axes, scales, and what the plotted values represent should be described in the caption or in §5.3.","section":"Fig. 4"},{"comment":"The phrase 'the Meta Quest 3 is an AR-capable headset featuring color passthrough' is potentially confusing because the device is also marketed as a VR headset; consider using 'mixed reality' or 'passthrough-based AR' for precision.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"This is a modest exploratory study whose central quantitative claims are conditional on a Wizard-of-Oz simulation. The most important fix is to align the abstract and conclusions with what was actually evaluated—an AR visualization interface with simulated handwriting parsing—and to temper the statistical language. With those changes, the paper could be acceptable as an extended-abstract contribution; without them, the headline claims overstate the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a small, honest exploratory study of an AR headset for 3D math surfaces with handwriting-style input, and the handwriting recognition is done by a human operator behind a loading spinner. The paper has a real prototype, clear writing, and a candid limitations section. But the abstract sells the Wizard-of-Oz condition as a working OCR pipeline, and the engagement and ease-of-use measurements were all taken under that simulated condition. So the real contribution is a design-space exploration, not a validated input method.\n\nWhat is actually new: combining HMD AR with handwriting-like equation input for 3D surfaces. The cited Augmented Math work does OCR on mobile for 2D plots, and GeoGebra requires manual equation entry; neither offers this combination. The 3D function plotter and graph manipulation are implemented, and the comparison against a whiteboard and two GeoGebra variants is a reasonable baseline. It is believable that a headset with grab-and-move interaction is more engaging than a whiteboard. The authors also deserve credit for explicitly saying in Sec. 3.3 that they simulated the recognizer and in Sec. 7 that operator variability may have affected perceived ease of use.\n\nNow the soft spots, in proportion. The load-bearing assumption is that the WoZ operator is a fair stand-in for real OCR. Participants never experienced a misread, never corrected an equation, and the latency was operator-dependent. That means the 4.6/5 rating for handwritten-equation input and the ease-of-use parity with GeoGebra Desktop are conditional on a perfect recognizer. A real OCR system will misread, require correction loops, and add variable delay, all of which could erode the measured advantages. The abstract says \"real-time handwritten equation parsing\" and \"most effective in aiding problem-solving\" without these caveats. The effectiveness result is also a self-reported ranking, not observed problem-solving performance, and the paired t-tests on 10 participants with no multiple-comparison correction are thin. The sample is homogeneous and the problems were not graded.\n\nNone of this is fatal for a workshop-format exploration. The paper is what it says in the introduction: a Wizard-of-Oz prototype. But the claims need to be reframed to match the evidence. For this extended abstract, the authors should at least adjust the title and abstract to disclose the simulation up front and release the questionnaire items. A serious referee could help them do that.\n\nWho this is for: HCI people working on AR education and anyone interested in how Wizard-of-Oz prototyping interacts with user studies. I would not cite it in my own work without a real OCR follow-up, but I would send it to peer review rather than desk reject — the prototype is useful, the limitations are acknowledged, and the study can be improved with relatively modest changes.","headline":"A genuinely exploratory AR prototype with an honest limitations section, but the Wizard-of-Oz parser means the headline claim of 'handwritten equation parsing' is not actually tested.","tokens_in":9086,"tokens_out":2149,"would_cite":false,"duration_ms":23733,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An AR headset that accepts handwritten equations and plots their 3D surfaces was rated more engaging than desktop and mobile graphing tools, while matching the desktop tool's ease of use.","keywords":["augmented reality","handwritten input","equation parsing","3D function visualization","head-mounted display","Wizard-of-Oz","mathematics education","user study"],"falsifier":"Rerun the study with an actual handwriting recognizer running on the headset instead of the human operator; if engagement or ease-of-use ratings drop to the level of GeoGebra AR or GeoGebra Desktop, or if users report frustration with recognition latency and errors, the central claim would be refuted.","tokens_in":8007,"feed_emoji":"🥽","tokens_out":7901,"duration_ms":77477,"temperature":0.7,"pith_summary":"This paper claims that combining handwritten equation input with real-time 3D surface rendering on an augmented-reality headset produces a more engaging and equally usable mathematics visualization tool than existing options. The authors built Breaking the Plane, a headset application that detects an equation written on a whiteboard, plots its 3D graph in the user's field of view, and lets the user grab, rotate, and rescale the graph. In a within-subjects study with 10 multivariable-calculus students, the system scored significantly higher on engagement than a plain whiteboard, GeoGebra AR, and GeoGebra Desktop; it matched GeoGebra Desktop on ease of use; and it was most often ranked as the most effective aid for the problem-solving questions. Users also reported a strong willingness to use the system again, with the handwriting input and direct graph manipulation cited as the reasons. The paper presents this as evidence that head-mounted AR with handwriting-style input can make mathematical visualization both more engaging and no harder to use than established desktop tools, while acknowledging that the parsing was simulated by a human operator.","feed_headline":"AR that plots your handwritten equations wins on engagement","feed_subtitle":"In a 10-person study it matched GeoGebra Desktop for ease of use and ranked first for problem-solving.","key_machinery":"The load-bearing mechanism is the input loop: a user writes an equation on a whiteboard; an operator watching the headset's camera feed types it into the system over a WebSocket, with an \"OCR Processing...\" indicator masking the delay; a custom plotter evaluates the string expression and generates a procedural 3D mesh using the NCalc expression-evaluation library; and controller-based gestures let the user grab, rotate, scale, and pan the graph in the passthrough environment. The paper attributes the engagement and ease-of-use results primarily to this loop, because handwritten input accepts any notation a user already knows and frees the user from learning app-specific syntax.","core_discovery":"The paper introduces Breaking the Plane, an AR headset application that turns a handwritten two-variable equation into an interactive 3D surface plot rendered in the user's physical space. Its central finding is that this combination of handwritten input and head-mounted AR visualization was rated significantly more engaging than a plain whiteboard, GeoGebra AR, and GeoGebra Desktop, while being perceived as comparable in ease of use to GeoGebra Desktop. Participants ranked the system most often as the most effective tool for aiding problem-solving, and reported a high likelihood of future use. The authors interpret these results as evidence that HMD AR with handwriting-style input can support in-situ exploration of multivariable functions and remove the syntax barrier associated with existing graphing tools. The paper is careful to note that the parsing pipeline was simulated by a Wizard-of-Oz operator rather than an automatic recognizer.","pith_inferences":["I infer that part of the engagement gain may be a novelty effect of the headset plus the operator's near-perfect transcription; a production recognizer with visible latency and errors could narrow the gap.","I infer that the handwriting-input mechanism could extend to other notation-heavy fields such as physics or engineering, where students sketch equations and diagrams that a recognizer could turn into manipulable 3D objects.","I infer that a longitudinal or between-subjects replication would test whether the engagement advantage persists after users become accustomed to the headset and the novelty fades.","I infer that the comparable ease-of-use result suggests input method, not display technology, is the main usability bottleneck for AR math tools; improving recognition may matter more than further polishing graph rendering."],"forward_implications":["A head-mounted AR graphing tool can be as easy to use as a mature desktop application while being significantly more engaging, suggesting that AR need not trade usability for immersion.","Handwritten, syntax-free equation input removes the keystroke-learning barrier that prior work linked to low confidence with desktop graphing software.","Because participants ranked the system first or second for aiding problem-solving, a working automatic recognizer could make AR graphing a practical study aid for multivariable calculus.","The engagement advantage over a mobile AR app points to head-mounted, hands-free interaction, rather than AR itself, as the driver of the benefit.","The strong willingness-to-reuse scores suggest feasible adoption in educational settings if a production-ready handwriting recognizer is integrated."],"supporting_citations":[{"why":"Supplies the two baseline comparison systems, GeoGebra Desktop and GeoGebra AR, used in the within-subjects evaluation.","marker":"[2]"},{"why":"The prior AR system with OCR equation parsing and textbook-augmented visualizations that this paper extends to handwritten input and 3D surfaces.","marker":"[4]"},{"why":"Prior evaluation of GeoGebra AR showing it was practical but limited by small phone screens, motivating the head-mounted display form factor.","marker":"[10]"},{"why":"Reports students' low confidence with GeoGebra Desktop due to keystroke syntax, motivating the handwritten input approach.","marker":"[13]"},{"why":"Systematic review evidence that AR in mathematics education increases motivation and learning, framing the engagement hypothesis.","marker":"[1]"},{"why":"Teacher survey identifying the desire for students to create AR content as a key adoption factor, motivating real-time equation authoring.","marker":"[5]"},{"why":"Systematic review listing the difficulty of developing AR materials as a disadvantage, motivating automatic equation parsing.","marker":"[11]"}],"fun_headline_variants":["Handwriting becomes 3D plots in AR, study finds","AR with handwritten input scores top marks for engagement","Your scribbles, rendered in 3D: AR wins engagement test","AR visualizes your equations in 3D, out-engages rivals","Handwritten math meets AR: more engaging than whiteboard"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that a human operator typing the user's handwritten equations into the system feels to the user like real automatic handwriting recognition, so that the measured engagement and ease-of-use ratings would carry over to a production OCR system.","fun_headline_variants_meta":{"raw":{"variants":["Handwriting becomes 3D plots in AR, study finds","AR with handwritten input scores top marks for engagement","Your scribbles, rendered in 3D: AR wins engagement test","AR visualizes your equations in 3D, out-engages rivals","Handwritten math meets AR: more engaging than whiteboard"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000343,"raw_usage":{"total_tokens":1852,"prompt_tokens":875,"completion_tokens":977,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":890}},"tokens_in":491,"tokens_out":977,"duration_ms":10719,"temperature":1.0,"reasoning_tokens":890,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:21:03.217867+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the study with an actual handwriting recognizer running on the headset instead of the human operator; if engagement or ease-of-use ratings drop to the level of GeoGebra AR or GeoGebra Desktop, or if users report frustration with recognition latency and errors, the central claim would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the two baseline comparison systems, GeoGebra Desktop and GeoGebra AR, used in the within-subjects evaluation."},{"cited_title":"Augmented Math: Authoring AR-Based Explorable Explanations by Augmenting Static Math Textbooks","cited_arxiv_id":"2307.16112","evidence_quote":"The prior AR system with OCR equation parsing and textbook-augmented visualizations that this paper extends to handwritten input and 3D surfaces."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior evaluation of GeoGebra AR showing it was practical but limited by small phone screens, motivating the head-mounted display form factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports students' low confidence with GeoGebra Desktop due to keystroke syntax, motivating the handwritten input approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Systematic review evidence that AR in mathematics education increases motivation and learning, framing the engagement hypothesis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Teacher survey identifying the desire for students to create AR content as a key adoption factor, motivating real-time equation authoring."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Systematic review listing the difficulty of developing AR materials as a disadvantage, motivating automatic equation parsing."}],"review_version":1}