{"id":"26b663d1-b27e-4041-9cfe-bcbae15c167f","arxiv_id":"2411.15535","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A robot-integrated projection system for teaching shortest path algorithms was built, and two small preference studies found most students favored the robot-based demonstration over the on-screen animation.","lead":"The paper presents Timmy, a GoPiGo robot with overlaid projections that demonstrates Dijkstra, A*, and Bellman-Ford shortest path algorithms by moving across a user-drawn graph. Two small user studies (n=10, n=6) found that most participants preferred the robot-synced demonstration over a screen-only animation, though usability was rated higher for the screen.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pilot preference finding is contradicted by the paper's own Discussion: Section 4.2.2 reports 4/6 preferred robot, while Section 5 says pilot participants preferred the screen-based learning condition.","rationale":"I read the paper as a preliminary, honestly scoped workshop study: the claim is about preference and engagement, not measured learning, and the authors say so explicitly. The initial study's unanimous robot preference is genuine evidence that a passively observed robot-synced demonstration can be engaging. The reader's handling confound is a legitimate external-validity concern: in the pilot, the robot condition required manual placement, which may explain lower SUS and Q4 results. However, the most load-bearing problem is internal and prior to interpretation: the pilot's headline preference is stated one way in Section 4.2.2 and the opposite way in Section 5, or else a result central to the teaching claim is omitted. This is not a matter of disagreeing with the consensus or asking for more data; it is an inconsistency in the reported facts that the central claim relies on. The proposed audit can settle it from existing records, so the paper is salvageable if the contradiction is a summary or wording error. I therefore retain the reader's CONDITIONAL verdict rather than escalate to reject. I only partially agree with the reader's weakest_assumption: the handling confound is real, but it assumes the pilot preference data have already been read correctly, and that is exactly what the Section 5 contradiction calls into question.","tokens_in":6279,"tokens_out":9866,"duration_ms":84106,"concrete_test":"Audit the raw pilot records: (1) confirm the number of participants and the counterbalancing order; (2) list each participant's answer to the Q1 preference question; (3) check whether a 'better for learning' question was asked and, if so, record those answers. Then map Section 4.2.2's '4/6 preferred the robot' and Section 5's 'preferred the screen-based learning condition' onto the corresponding questions. If both refer to Q1, one statement must be corrected. If they refer to different questions, the missing question must be reported and the abstract should be adjusted to say the robot was preferred for engagement but the screen was preferred for learning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the pilot preference data being reported consistently. They are not. Section 4.2.2 reports: 'Four out of six (66%) participants preferred the robot-synced over the on-screen-only condition.' Section 5 then states: 'in the pilot study, where participants had to build their own graphs and interact with the robot, they preferred the screen-based learning condition.' On their face, these describe opposite outcomes of the same pilot. If the Section 5 sentence refers to a separate, unreported question about which condition is better for learning, then the paper has omitted the result most relevant to its teaching claim while featuring the engagement preference in the abstract. If it refers to the same Q1 question, one of the two statements is wrong. Either way, the abstract's 4/6 majority is not an unambiguous fact: the paper simultaneously tells the reader that the robot was preferred and that the screen was preferred for learning. The pilot's usability results (screen SUS 80.42 vs robot 72.92; Q4 ease-of-handling 4/6 screen) suggest the Discussion may have collapsed 'easier to handle' into 'preferred learning condition,' but the text does not say so. This is more load-bearing than the handling confound because it concerns what the data mean even before confounds are considered.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents Timmy, a GoPiGo3 robot augmented with a ceiling-mounted projector that renders a graph-drawing application around the robot, with the robot physically executing the steps of Dijkstra, A*, and Bellman-Ford shortest path algorithms in synchrony with the visualization. The system is evaluated through two within-subject preference studies: an initial observation-only study (n = 10, Dijkstra only) and an interactive pilot (n = 6, all three algorithms), comparing a robot-synced mode against an on-screen-only mode using NASA-TLX (initial study) and SUS (pilot) questionnaires plus preference questions. The paper reports that all 10 initial participants preferred the robot-synced mode and perceived it as better for learning, while in the pilot 4/6 preferred the robot-synced mode as the liked condition, all 6 chose it for teaching primary school children, and 4/6 found the screen easier to handle. The authors explicitly state that both studies measured preferences, not learning outcomes, and they conclude with calls for larger studies and methodological refinements, also listing technical limitations of the GoPiGo platform.","tokens_in":6495,"tokens_out":13798,"duration_ms":108623,"significance":"Taken at face value, the preference data provide modest self-report evidence that embodied, physically traversing algorithm demonstrations can be engaging and attention-maintaining for university students, and that users perceive such demonstrations as suitable for teaching younger children. The paper's concrete strengths are the replicable system description (grid-normalized movement calibration, file-based command synchronization between the JavaScript application and the robot, and the select/poke/traverse/celebrate action vocabulary), the transparent per-participant workload reporting in Figure 3, and the honest, explicit acknowledgement that no learning outcome was measured and that sample sizes are small. Because no derivation, fitting, or parameter estimation is involved, circularity is not at issue; the self-report data directly support the descriptive claims. The significance is limited by the inconsistent reporting of the pilot result (Section 4.2.2 versus Section 5), the confound of extra manual handling in the robot condition, and the absence of any inferential statistics.","major_comments":[{"comment":"The pilot preference finding is stated inconsistently. Section 4.2.2 reports that 'Four out of six (66%) participants preferred the robot-synced over the on-screen-only condition,' while Section 5 states that 'in the pilot study, where participants had to build their own graphs and interact with the robot, they preferred the screen-based learning condition.' These two sentences cannot describe the same preference question. If the Section 5 sentence refers to a question about which condition is better for learning (an analogue of the initial study's Q2), then the paper has omitted the result most relevant to its teaching claim while featuring the engagement preference in the abstract; if it refers to the preference question reported in Section 4.2.2, then one of the statements is wrong. The Discussion's summary that participants' preference 'shifted based on the task' presupposes that this contradiction has been resolved. The authors must report the pilot questionnaire item by item, give the counts for each item (including any 'better for learning' question), and rewrite Section 5 so it is consistent with Section 4.2.2. Because the abstract's 'initial findings' rest in part on the 4/6 pilot majority, this inconsistency is load-bearing and must be fixed before the central claim can be evaluated.","section":"§4.2.2 and §5"},{"comment":"The robot-versus-screen comparison is confounded in two ways that the preference conclusions do not control for. First, in the robot-synced condition of the pilot, participants had to manually position the robot before each algorithm run ('users were first given additional instructions on where to place the robot... This process was repeated each time they chose an algorithm'), whereas the on-screen-only condition required only clicking a button, as participant P_P01's remark 'Just clicking a button' makes explicit. The SUS means (screen 80.42 versus robot 72.92) and the Q4 handling result (4/6 screen) are plausibly driven by this asymmetry, so the 4/6 Q1 preference for the robot cannot be unambiguously attributed to the delivery medium as opposed to the added handling burden. Second, the two conditions are not informationally equivalent: Section 2.2 states that in robot-synced mode 'edge colouring is omitted,' so the green edge-exploration animations shown on screen are absent from the robot condition; the robot's 'poke' actions are not argued to convey the same algorithmic detail. The authors should equalize handling effort across conditions, measure handling burden separately from visualization modality, and discuss the omitted edge coloring when interpreting the preference results.","section":"§4.1, §4.2.1, §2.2"},{"comment":"The paper's framing claims more than the measurements support. The title is 'Teaching Shortest Path Algorithms With a Robot and Overlaid Projections,' the conclusion states that 'results suggest that robots can maintain users' attention and show promise as educational tools,' and Section 5's summary about participants' preference shifting 'based on the task' invokes learning-related preference. Yet the abstract itself states that 'In both studies we investigated the preferences towards the system and not the teaching outcome,' and Section 5 concedes that 'future research should explore the effectiveness of using such teaching methods by testing participants' understanding of the content being taught.' No learning measure, knowledge test, or retention data appear anywhere in the manuscript. The authors should either add a learning-related outcome or systematically rephrase the title, abstract, and conclusion so that they claim only engagement and preference evidence; the initial study's all-10 'better for learning' response must be explicitly labeled as a subjective perception, not a measured learning gain.","section":"Title, §6, Abstract"}],"minor_comments":[{"comment":"The protocol sentence reads 'We recruited 6 participants, all university students (n = 10, aged 18 to 25)'; the parenthetical n = 10 contradicts the stated n = 6 and should be corrected.","section":"§4.1"},{"comment":"The phrase 'with n = 3 participants per condition, and the conditions were counterbalanced' is misleading for a within-subject design; it should say that three participants received each condition order.","section":"§4.1"},{"comment":"The preference-question numbering appears shifted relative to the questionnaire in Section 3.1: in the pilot, 'Q3' is the primary-school teaching question and 'Q4' is ease of handling, whereas in the initial study Q3 was ease of handling and Q4 was recommendation to colleagues; clarify the pilot questionnaire numbering explicitly.","section":"§4.2.2"},{"comment":"The closing sentence 'user preference for educational methods depends significantly on user preference' is tautological; rephrase to say that preferences varied across participants and depended on individual priorities such as speed versus engagement.","section":"§4.2.2"},{"comment":"No inferential statistics accompany any reported count or SUS mean; given n = 6, the '66%' framing is fragile, and reporting exact binomial confidence intervals or effect sizes would prevent over-reading; also, the NASA-TLX and SUS instruments are used without citations, which should be added.","section":"§4.2.2"},{"comment":"All pilot participants experienced the three algorithms in a fixed order (Dijkstra, then A*, then Bellman-Ford), so order effects cannot be excluded and should be acknowledged.","section":"§4.1"},{"comment":"Several small textual errors should be fixed: 'educators educators' and 'One such options' in Section 1, the corrupted token '/envel⌢pe-⌢pen' in the author-address line, and the stray 'Both-' in the initial study's Q5 option list.","section":"§1 and header"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short workshop paper (CEUR-WS, HCI Slovenia 2024) and should be judged at that scope. The main gate for acceptance is the resolution of the Section 4.2.2 versus Section 5 contradiction: if the authors can report the pilot questionnaire item by item and reconcile the statements, the paper is publishable as a feasibility report; if it turns out that the pilot actually exhibited a screen preference on the central question, the abstract and the 'engaging tool' claim would need substantial revision, and I would then regard rejection as appropriate. The sample sizes are too small for the quantitative claims to carry weight on their own, but the authors are properly hedged about this. I saw no novelty-disclosure or citation-integrity concerns; the related-work section is thin but acceptable for the venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on educational robotics, but read the pilot sections twice: Section 4.2.2 reports that four of six pilot participants preferred the robot-synced mode, while Section 5 states that in the pilot study 'they preferred the screen-based learning condition.' Those cannot both be true as written. The abstract leans on the pilot for its 'engaging tool' claim, so this is a load-bearing inconsistency, not a typo. If the Discussion is right, then the preference data are misreported; if the results section is right, then the Discussion misstates the outcome. Either way, the paper needs a correction before its findings can be trusted.\n\nWhat is actually new: the Timmy system itself — a GoPiGo robot synced with projected graph visualizations and a graph editor, supporting Dijkstra, A*, and Bellman-Ford. That integration is not in the cited prior GoPiGo work. The authors also report original preference data from two within-subject studies. To their credit, they are explicit that they measured preferences, not learning, and they openly list the robot's practical problems (battery, VNC, slow Pi). The initial study's unanimous preference for the robot is a clean, if small, descriptive result.\n\nSoft spots beyond the contradiction: the pilot protocol says '6 participants' and then immediately says '(n = 10, aged 18 to 25)' — a copy-paste error that does not inspire confidence. There are no inferential statistics, so the 4/6 preference is just a count. The robot condition required manual handling (positioning the robot before each run) that the screen condition did not, and the authors acknowledge this as a possible cause of the lower SUS score. No code or data are provided, which is common for a workshop paper but still limits reproducibility.\n\nThe system description is clear and the authors seem genuinely engaged with the limitations. The contradiction, though, means the paper is not coherent on its own terms right now. If the authors fix the reporting discrepancy, clarify what was actually asked in the pilot, and add the per-question numbers, this becomes a useful data point for the HCI/educational-robotics community. It deserves a serious referee, but not a clean accept as is.","headline":"A likeable small educational-robotics paper with an internal contradiction: the Discussion says pilot participants preferred the screen, while the results section says 4/6 preferred the robot.","tokens_in":7019,"tokens_out":2155,"would_cite":false,"duration_ms":20847,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A robot that physically walks through projected graphs was preferred by all ten observers in one study and by four of six active users in a pilot over screen-only animation, though usability ratings favored the screen.","keywords":["GoPiGo","shortest path algorithms","teaching","algorithms","graphs","robots","projections"],"falsifier":"A controlled experiment that removes the extra handling—an autonomous robot with pacing matched to the screen and no manual placement—and then measures learning with pre/post tests would settle the claim: if the robot-synced mode no longer wins preference or shows no learning improvement over screen-only, the central claim that the physical robot itself aids teaching would be refuted.","tokens_in":6101,"feed_emoji":"🤖","tokens_out":8854,"duration_ms":74001,"temperature":0.7,"pith_summary":"This paper asks whether a physical robot can make abstract graph algorithms—Dijkstra's, A*, and Bellman-Ford—tangible enough that learners prefer it over a plain screen animation. It presents Timmy, a GoPiGo robot that moves across a projected graph while a JavaScript application colors edges and shows pseudocode alongside distance and predecessor updates. Two small within-subject studies compared the robot-synced mode with an on-screen-only mode: in the observation-only study (n=10) all participants preferred the robot, and in the interactive pilot (n=6) four of six did, with everyone saying they would recommend both. Self-reported cognitive load stayed manageable, around 20 percent on the NASA-TLX workload scale, and both modes received above-average System Usability Scale scores. The paper's conclusion is cautious: robots are engaging and hold attention, but preference depends on how much handling the robot requires, and teaching effectiveness still needs larger studies.","feed_headline":"A robot that walks graph algorithms wins student preference","feed_subtitle":"Most learners preferred watching a robot traverse graphs to screen-only animation, but less so when they had to handle it.","key_machinery":"The central object is Timmy, a GoPiGo robot paired with a top-down projector and a JavaScript graph-drawing application. The projector displays the graph, pseudocode, and distance and predecessor updates on the surface around the robot, and in robot-synced mode Timmy's physical movements replace the green animated edge traversal of the screen-only mode. Synchronization runs through text files containing vertex coordinates and adjacency lists, a watcher script that stores the robot's position and orientation in JSON, and timing calibrated from measured movement and turn rates. Timmy executes four primitives—select a vertex, poke a vertex, traverse the shortest path, and celebrate—so a learner watches a physical agent act out the algorithm's decisions. This mechanism is what translates an abstract computation into visible, embodied motion, the feature the paper hypothesizes will improve engagement.","core_discovery":"On the paper's own terms, the discovery is a proof of concept: a synchronized robot-plus-projection environment can demonstrate shortest path algorithms in a way that learners find engaging and mostly prefer to screen-only animation. Concretely, all ten passive observers chose the robot as their preferred condition and believed it was more effective for demonstrating the algorithms, while four of six interactive pilot participants preferred it. The paper also reports that preference shifts when participants have to position the robot themselves, that the screen condition received a higher mean usability score in the pilot, and that all pilot participants recommended using both modes and would welcome robots in class. The authors frame this as evidence that robots offer an engaging tool for teaching advanced algorithmic concepts, while cautioning that they measured preference, not learning outcomes.","pith_inferences":["Beyond the paper, if manual handling is what lowered the robot's preference in the pilot, an autonomous robot that positions itself should restore the near-unanimous preference seen in the passive study; this could be tested by comparing self-moving and manually placed conditions.","Beyond the paper, combining the robot's physical traversal with on-screen edge coloring would address both pilot complaints about the robot being slow and hard to track, since the screen would keep a persistent visual record while the robot adds embodiment.","Beyond the paper, the robot's slower pace, which one participant saw as a drawback, might become a teaching advantage if future research shows that slower, embodied demonstrations improve retention; a pre/post test of learning would settle that.","Beyond the paper, the split between liking the robot and finding it harder to handle suggests that engagement and usability are separable dimensions, so educational-robot evaluations should measure both rather than treat preference as a single score."],"forward_implications":["If the findings hold, robot-synced demonstrations can serve as an engaging supplement to screen animations, but not as an automatic replacement: the screen was rated more usable in the pilot, where users had to handle the robot.","The unanimous preference for the robot in the passive observation study points to demonstration-style teaching, with the instructor operating the robot, as the most promising near-term classroom use.","All pilot participants said they would recommend both modes and would like robots in class, implying that offering both interfaces can accommodate different learning styles and preferences.","The pilot's unanimous view that the robot would be better for teaching primary-school children marks a natural target audience, although the authors note that this does not establish suitability for advanced concepts.","Because both studies captured preferences rather than learning gains, the direct corollary is that engagement must still be shown to translate into understanding before classroom adoption is justified."],"supporting_citations":[{"why":"Establishes robots in the classroom as an educational tool, motivating the research direction.","marker":"[4]"},{"why":"Supplies the GoPiGo robot platform that Timmy is built on.","marker":"[5]"},{"why":"Defines the shortest path problem family the system teaches.","marker":"[9]"},{"why":"Provides Dijkstra's algorithm, one of the three demonstrated algorithms.","marker":"[10]"},{"why":"Provides A*, one of the three demonstrated algorithms.","marker":"[11]"},{"why":"Provides Bellman-Ford, one of the three demonstrated algorithms.","marker":"[12]"},{"why":"Supports the pilot participants' view that robots promote learning for primary school children.","marker":"[13]"},{"why":"Supplies prior evidence on measuring robots' effectiveness in teaching computer science, framing the evaluation.","marker":"[14]"}],"fun_headline_variants":["Robot teaches shortest paths, wins student preference","Robot demo of graph algorithms beats screen-only","Learners prefer robot-led algorithm lessons","Robot + projection: engaging way to teach algorithms","Proof of concept: robot makes algorithms tangible"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparisons assume that differences in preference come from the robot's physical visualization, even though the robot condition also required extra manual positioning, ran more slowly, and had connection problems—so the preference could stem from novelty or added effort rather than from the value of the physical movement itself.","fun_headline_variants_meta":{"raw":{"variants":["Robot teaches shortest paths, wins student preference","Robot demo of graph algorithms beats screen-only","Learners prefer robot-led algorithm lessons","Robot + projection: engaging way to teach algorithms","Proof of concept: robot makes algorithms tangible"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000148,"raw_usage":{"total_tokens":1155,"prompt_tokens":879,"completion_tokens":276,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":210}},"tokens_in":495,"tokens_out":276,"duration_ms":3025,"temperature":1.0,"reasoning_tokens":210,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:10:04.572518+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment that removes the extra handling—an autonomous robot with pacing matched to the screen and no manual placement—and then measures learning with pre/post tests would settle the claim: if the robot-synced mode no longer wins preference or shows no learning improvement over screen-only, the central claim that the physical robot itself aids teaching would be refuted.","supporting_citations":[{"cited_title":"Bellman, On a routing problem, Quarterly of applied mathematics 16 (1958) 87–90","cited_arxiv_id":null,"evidence_quote":"Provides Bellman-Ford, one of the three demonstrated algorithms."},{"cited_title":"Baxter, E","cited_arxiv_id":null,"evidence_quote":"Supports the pilot participants' view that robots promote learning for primary school children."},{"cited_title":"Fagin, L","cited_arxiv_id":null,"evidence_quote":"Supplies prior evidence on measuring robots' effectiveness in teaching computer science, framing the evaluation."},{"cited_title":"Gallo, S","cited_arxiv_id":null,"evidence_quote":"Defines the shortest path problem family the system teaches."},{"cited_title":"Reich-Stiebert, F","cited_arxiv_id":null,"evidence_quote":"Establishes robots in the classroom as an educational tool, motivating the research direction."},{"cited_title":"Industries, Gopigo: An educational robot for learning programming and robotics, https://www.dexterindustries.com/gopigo, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the GoPiGo robot platform that Timmy is built on."}],"review_version":1}