{"id":"222a8a88-0c5b-492c-b581-aac1e0b8dcaf","arxiv_id":"2412.16892","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A design toolkit for AI-plus-mixed-reality physical task guidance: six design considerations, 36 design patterns, and an interaction canvas, derived from student prototypes and tested in a small user study.","lead":"This paper introduces MixITS-Kit, a set of design tools for building AI-powered mixed reality systems that guide people through physical tasks. The toolkit was derived from eight student-designed prototypes and evaluated in a small user study with eight participants.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim hinges on patterns elicited from eight student prototypes in one author-taught course; no evidence yet shows they transfer to professional MixITS design. A coverage study of documented real-world MixITS systems would settle whether the catalog is complete.","rationale":"The reader's weakest_assumption identifies exactly the transferability/external-validity gap that I see as load-bearing. The paper is transparent, the derivation is documented, and the toolkit is plausible; the concern does not invalidate the contribution but warrants keeping the verdict CONDITIONAL. The participant-count discrepancy (abstract says nine, §6 says eight, yet results quote P9/P10) is a secondary correctness issue that should be corrected but is not central. I would not move the verdict.","tokens_in":25158,"tokens_out":4518,"duration_ms":46051,"concrete_test":"Select 5–10 documented MixITS systems from §2.2 and related work (e.g., Feiner et al. 1993, YouMove, ARTiST, SIGMA, Skillab, ARDW). For each, have two independent coders decompose the system's design decisions and interaction features into episodes, then attempt to map each episode onto one of the 36 patterns and one of the eight gulfs, with a pre-registered coverage threshold (e.g., ≥80% of episodes map without forced re-labeling). If coverage falls below threshold, the catalog is incomplete or curriculum-biased and the practitioner-oriented claim must be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is made in §3 and acknowledged in §8.1: the 36 design patterns and six considerations were elicited from eight low-fidelity prototypes built by 25 graduate students in one 10-week course taught by the authors. The central claim in §9 — that MixITS-Kit 'offers a structured set of tools that fulfill gaps' for practitioners — therefore depends on the classroom artifacts being representative of professional MixITS design. The only evaluation (§6) uses eight student participants with mean self-rated MixITS expertise 2.0/5, no baseline condition, and tasks authored by the same team around the very gulfs used to build the catalog; Task 2 additionally asks participants to recognize labels from the catalog, which is partly circular. If the course materials, assignments, readings, and in-class role-play shaped the observed problems, the catalog may encode curricular constraints rather than recurring MixITS design problems. The paper itself concedes in §8.1 that 'future work should expand the MixITS-Kit with data collected from more experienced professionals.' This does not make the work internally inconsistent, but it is the load-bearing weak point for the practitioner-directed claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MixITS-Kit, a design toolkit for Mixed Reality Intelligent Task Support (MixITS) systems, comprising six high-level design considerations, 36 design patterns grouped under eight interaction gulfs, and an Interaction Canvas. The toolkit is derived through reflexive thematic analysis and design-pattern elicitation on eight low-fidelity prototypes built by 25 graduate students in a 10-week author-taught course. The authors evaluate MixITS-Kit in an asynchronous take-home study with eight participants, who used the toolkit to analyze and solve design problems for a fictional rock-climbing guidance app, and report generally positive self-ratings and moderate pattern-recognition accuracy. The paper concludes that the toolkit can serve as a valuable resource for practitioners and researchers.","tokens_in":25290,"tokens_out":7190,"duration_ms":63472,"significance":"If the elicited content transfers beyond the classroom setting, MixITS-Kit addresses a real gap: the existing design guidance covers AI and MR separately or only partly at their intersection, while little structured support exists for the combined design space of physical task guidance. The manuscript's strengths include the concrete and inspectable artifact in Table 1 (the pattern catalog), the anchoring of the analysis in Norman's Gulfs of Execution and Evaluation, the use of an established evaluation framework for HCI toolkits (Ledo et al.), and a transparent acknowledgment of the main limitations, including curricular bias and the need for data from experienced professionals. The central risk is that the practitioner-directed claim rests on an unverified assumption about the representativeness of eight student-team prototypes, and the evaluation does not provide strong evidence about transferability or comparative value. The contribution is best characterized as an initial, structured resource that needs further validation before broad practitioner guidance claims are warranted.","major_comments":[{"comment":"The MixITS-Kit is elicited exclusively from eight low-fidelity prototypes developed by 25 graduate students in a single 10-week course taught by the authors (§3, §4.2), and §8.1 acknowledges that the teaching approach and selected materials may have introduced biases. The abstract's claim that the toolkit 'can serve as a valuable resource for practitioners' and §9's claim that it 'offers a structured set of tools that fulfill gaps in existing design aids' depend on these classroom artifacts being representative of real-world MixITS design problems. No evidence is presented for such transferability, e.g., a coverage analysis of the published MixITS systems cited in §2.2 or design data from professional practitioners. Since §8.1 explicitly defers the collection of professional data to future work, the central claim currently exceeds the evidence; the paper should either add an external-validity check or temper the abstract and conclusion to describe a catalog for early-career designers with its external validity still open.","section":"§3, §8.1, §9"},{"comment":"In Task 2, each participant's 'correct' identification of a peer's design pattern is scored against the authors' own labeling (the 'originally intended ones,' §6.4). Because the solutions being classified were produced with the same catalog and the reference labels come from the authors' Table 1, the exact-match rate of half of the participants measures agreement with the authors' interpretation rather than an independent demonstration of shared vocabulary. In addition, the study has no baseline condition, so the favorable self-reports in Figure 6 (e.g., 'It would take me longer to solve the task without the toolkit,' M=4.5) cannot be compared against working without the toolkit or against using the underlying AI/MR guideline sets separately. The paper should either add a comparison condition or explicitly restrict the claim to perceived usefulness in a single session with novice designers.","section":"§6.4–§6.6"},{"comment":"The results section reports that in Task 1 two of eight participants applied patterns from non-anticipated gulfs with justifications that did not align with the gulf concepts, and in Task 2 only four of eight identified the exact design pattern and five of eight identified the correct gulf. Section 7.1 nonetheless concludes there was a 'high level of performance' and calls the results promising. With a 50% exact-match rate and no baseline, the evidence is more naturally read as showing that the toolkit is learnable but that pattern recognition is unreliable; the interpretation should be calibrated to these numbers.","section":"§6.6, §7.1"}],"minor_comments":[{"comment":"The recognition counts in §6.6 are reported in a way that does not add up for the eight participants: 'Four participants successfully recognized the correct design pattern, while five identified the correct gulf. One participant misidentified the pattern, and three incorrectly identified the gulf.' Please clarify whether these categories are mutually exclusive and how the remaining participants (if any) are classified, or report full contingency counts.","section":"§6.4–§6.6"},{"comment":"There are typographical and phrasing issues in the evaluation section: 'We instructed participants to to work directly' (§6.3), 'experinece' (§6.1), and 'we consider this participants are representative' (§6.1). These should be corrected.","section":"§6.1, §6.3"},{"comment":"Reference [5] appears garbled ('BURTON R. R. and ITS. International Conference. 1982. Diagnosing Bugs in a Simple Procedural Skill, Sleeman.'); please fix the citation entry.","section":"References"},{"comment":"The statement in §6.6 that 'Participants identified the most challenging aspects of learning the canvas and design patterns' does not indicate how these challenges were measured or aggregated; please clarify whether these are open-ended responses or quantitative items.","section":"§6.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable HCI toolkit contribution, and the pattern catalog itself is a useful artifact. My main concern is scope: the practitioner-directed claims rest on student-generated prototypes, and the evaluation has no baseline. A revision that substantially tempers the claims, or adds an external-validity analysis, would make the contribution publishable. Given the central claim is defensible but the evidence needs strengthening, I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Genuinely new here is the eight-gulf extension of Norman's model to the human/AI/environment triangle, the 36-pattern catalog, and the Interaction Canvas. The patterns are concrete, the gulf taxonomy gives designers a clean way to map MixITS problems, and the authors are honest about the biggest limitation: the data come from eight student teams in one 10-week course taught by the authors. That honesty is in Section 8.1, not buried.\n\nWhat the paper does well: the pattern table is well-organized and the examples are specific enough to act on. The design considerations (teaching vs. directing, interaction timing, error handling, evolving context, building trust) read as grounded in the eight prototypes rather than generic AI/MR advice. The evaluation uses Ledo et al.'s toolkit research goals and reports medians plus MAD, which is appropriate for a small sample. Task 2 is a clever idea—asking participants to recognize patterns in each other's solutions is a reasonable proxy for shared vocabulary.\n\nSoft spots, in proportion. Transferability is the real issue: patterns were elicited from graduate student prototypes in an author-taught course, and the evaluation used eight students whose self-rated MixITS expertise averaged 2.0/5. There is no baseline, no comparison with expert designers, and no objective measure of design quality. In Task 1, 'success' is judged against the authors' own gulf/pattern labels; in Task 2, correct recognition is scored against the authors' original labeling. So the shared-vocabulary evidence is partly circular. The authors concede this in Section 8.1, but the concession doesn't remove the problem for the practitioner-directed claim.\n\nThere are also internal inconsistencies a referee should catch: the abstract and introduction say nine participants, Section 6.1 says eight, and the results quote P9 and P10. The Task 2 numbers are murky—four recognized the pattern, one misidentified, three unaccounted for. These are small fixes, but they are real.\n\nWho this is for: HCI researchers working on AI+MR design toolkits, and practitioners who want a structured starting point for MixITS design. The central claim is modest enough that the evidence supports it, if you read 'resource' as 'helpful starting point for novices' rather than 'validated professional toolkit.' A coverage study of documented real-world MixITS systems would go a long way toward settling the transferability question.\n\nRecommendation: this deserves a serious referee. The framework and catalog are useful enough to justify revision; the evaluation needs strengthening, but the core contribution is real.","headline":"A genuinely useful eight-gulf framework and 36-pattern catalog for MixITS design, backed by honest but thin evidence from one student course; worth refereeing, not desk rejecting.","tokens_in":25902,"tokens_out":3194,"would_cite":true,"duration_ms":29293,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes MixITS-Kit, an interaction design toolkit that distills eight low-fidelity AI+MR task-guidance prototypes into six design considerations, 36 design patterns, and an Interaction Canvas.","keywords":["mixits","physical task guidance","mixed reality","human-AI interaction","interaction design toolkit","design patterns","interaction canvas","low-fidelity prototyping"],"falsifier":"An observational study of independently built MixITS systems during user testing would settle the catalog's completeness: if a substantial share of interaction breakdowns cannot be assigned to any of the 36 pattern labels, the claim that these are the recurring design problems of the domain is falsified.","tokens_in":24868,"feed_emoji":"🧩","tokens_out":9811,"duration_ms":84010,"temperature":0.7,"pith_summary":"Physical skills often need live instruction and feedback, but a human expert is not always available. The paper argues that systems combining artificial intelligence with mixed reality can fill that role, and that designing such systems is hard in a specific, discoverable way: problems arise at the junctions between user, AI, and the physical environment. To capture those problems, it proposes MixITS-Kit, a toolkit made of six high-level design considerations, 36 lower-level design patterns with example solutions, and an Interaction Canvas for analyzing interaction breakdowns. The toolkit was derived from eight low-fidelity prototypes built by graduate teams over a ten-week course, and a small evaluation suggests that newcomers can use it to analyze a design problem and generate concrete solutions. The paper's central claim is that these structured tools fill a gap left by design guidance that treats AI or mixed reality separately.","feed_headline":"36 patterns and a canvas for AI+MR task coaching","feed_subtitle":"A design toolkit built from eight student prototypes helps novices craft physical-task guidance with AI and mixed reality.","key_machinery":"The load-bearing mechanism is the Interaction Canvas built on the Gulfs of Execution and Evaluation model, which measures the gap between a user's intentions and the actions a system supports, and the gap between the system's displayed state and the user's interpretation. The model is applied not to a single user and system but to three entities: the human user, the AI-MR system, and the real environment. This yields eight named gulfs—human execution and evaluation toward AI and environment, and AI execution and evaluation toward human and environment—and each design pattern is assigned to one of these gulfs. The canvas asks the designer to specify, for a given interaction, who the actor is, what the target is, whether the actor achieved the goal, and whether the actor correctly interpreted the target's feedback. That same labeling process was used to sort the 63 functionalities extracted from the eight prototypes into the final 36 patterns, which gives the catalog its structure and provides designers with a shared vocabulary for naming breakdowns.","core_discovery":"On the paper's own terms, the central discovery is that the design problems of Mixed Reality Intelligent Task Support (MixITS)—AI-driven instruction and feedback for tasks in the real world—can be elicited and organized into a reusable interaction design toolkit. The authors claim that MixITS-Kit offers a structured set of tools that fulfill gaps in existing design aids that consider only AI or MR alone and tackle the unique challenges of task support situated in the real environment at different levels of abstraction. The toolkit consists of six design considerations (Teaching and Directing, Interaction Timing, Error Handling, Sensors and Actuators, Evolving Context, and Building Trust), 36 design patterns organized by eight interaction gulfs, and an Interaction Canvas that guides designers through execution and evaluation questions for each pair of actors. In an evaluation, eight participants with mixed AI and MR experience used the toolkit to diagnose fictional user issues, match them to patterns, and revise their solutions; most reported that the toolkit was learnable, supported creative solutions, and offered a shared vocabulary, while two of eight initially chose a pattern from a different gulf than expected.","pith_inferences":["The eight-gulf scheme is essentially an interaction graph with human, AI, and environment as nodes; that structure could be turned into a generative design probe or automated checklist for other situated AI systems, which the paper only gestures at in its future-work section.","The paper deliberately omits pattern frequencies, so the catalog currently cannot tell a novice which problems are most common; collecting frequency data from deployed systems would be a natural next dataset and would turn the taxonomy into an evidence-ranked resource.","Because the elicitation method relies on course artifacts, the most direct test of the toolkit's value is replication with professional practitioners; the paper asks for this explicitly, but it has not been done.","The paper speculates that taking the design considerations early reduces rework, but does not measure this; a longitudinal comparison of design iteration counts with and without the toolkit would put that claim on firmer ground."],"forward_implications":["Novice designers can move from a reported user problem to a concrete, sketched solution in about half an hour: median completion time was 27 minutes for the first task and 8.5 minutes for the revision task.","Because every pattern is labeled with a gulf, the toolkit gives design teams a shared vocabulary: in the peer-interpretation task, all eight participants correctly identified the actor and target of a solution, and half identified the exact pattern.","Designers are nudged to balance proactive and reactive AI interventions, preserve user agency when errors occur, and build trust through transparent explanations and error reporting, rather than treating guidance as one-way instruction.","The high-level considerations and low-level patterns are meant to be used together, so reflecting on a consideration such as Building Trust can revise and enrich an already proposed pattern solution.","The pattern catalog is explicitly a starting point, not a closed list, and the paper expects it to grow as more MixITS systems are built and studied."],"supporting_citations":[{"why":"Supplies the human-AI interaction guidelines that MixITS-Kit extends and positions against when adding MR and physical task context.","marker":"[3]"},{"why":"Provides the design-pattern elicitation method used to turn the eight prototypes into 36 problem-solution patterns.","marker":"[14, 67, 88]"},{"why":"Provides the reflexive thematic analysis procedure used to derive the six design considerations.","marker":"[15]"},{"why":"Supplies the toolkit research goals and evaluation strategy used in the take-home user study.","marker":"[47]"},{"why":"Supplies the Gulfs of Execution and Evaluation model that the Interaction Canvas extends to interactions among human, AI, and environment.","marker":"[59, 61]"},{"why":"Provides the explainable-AI-in-AR toolkit that the paper contrasts with, showing what a broader MixITS toolkit adds beyond explainability.","marker":"[90]"}],"fun_headline_variants":["36 patterns, 6 considerations, 1 canvas for AI+MR coaching","Design toolkit for physical task guidance with AI and MR","New toolkit helps design AI+MR coaching for physical skills","MixITS kit: 36 patterns and canvas for AI+MR task coaching","Fills the design gap between AI and MR for physical task guidance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the design challenges and solutions seen in eight low-fidelity student projects from one 10-week course taught by the same instructors also show up in real physical-task-guidance systems built by working professionals.","fun_headline_variants_meta":{"raw":{"variants":["36 patterns, 6 considerations, 1 canvas for AI+MR coaching","Design toolkit for physical task guidance with AI and MR","New toolkit helps design AI+MR coaching for physical skills","MixITS kit: 36 patterns and canvas for AI+MR task coaching","Fills the design gap between AI and MR for physical task guidance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000589,"raw_usage":{"total_tokens":2749,"prompt_tokens":915,"completion_tokens":1834,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":1744}},"tokens_in":531,"tokens_out":1834,"duration_ms":13000,"temperature":1.0,"reasoning_tokens":1744,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T06:00:23.719501+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An observational study of independently built MixITS systems during user testing would settle the catalog's completeness: if a substantial share of interaction breakdowns cannot be assigned to any of the 36 pattern labels, the claim that these are the recurring design problems of the domain is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the explainable-AI-in-AR toolkit that the paper contrasts with, showing what a broader MixITS toolkit adds beyond explainability."}],"review_version":1}