{"id":"a396202c-0dcb-45ba-b9da-c6f9b9661827","arxiv_id":"2608.11871","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"ATOM uses generative 3D and 2D geometry enhancement plus fingertip reachability and ergonomic scoring to automatically turn corners, edges, and surface patches on everyday objects into 0D, 1D, and 2D AR microgesture controls.","lead":"This paper presents ATOM, an AR system that turns everyday handheld objects like bottles, cartons, and mugs into touch controls by automatically finding small corners, edges, and flat surfaces on them. It is relevant to AR interaction design because it replaces manual setup with a pipeline that uses generative AI to clean noisy object geometry before detecting these interactive features.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Offline pre-scan and per-pose detection mean the user studies test cached elements, not runtime object-agnostic detection; the 'object-agnostic' claim is unsupported as stated.","rationale":"The reader's weakest_assumption identified the pre-scan requirement as one of four limitations, but framed it mainly as 'only works on the specific 10 known objects tested.' My stress-test sharpens this into the most load-bearing concern: even for those 10 objects, the user studies test only pre-cached detection results for fixed or suggested poses, so the headline 100% completion rate never exercises the actual detection pipeline under real-time, arbitrary-use conditions. The paper's own supplementary and Limitations sections confirm this (offline scanning, offline per-pose detection, non-real-time detection). This concern is load-bearing because the central innovation is 'object-agnostic' interaction; if the evidence only supports interaction with pre-scanned, pre-processed objects in controlled poses, the central claim is overstated. I do not think this overturns the paper: the authors explicitly aim 'towards object-agnostic' interaction, and the ablation comparisons within the cached-element setting remain relevant. The reader's CONDITIONAL verdict already accommodates this scope limitation, so the verdict should remain UNCHANGED. However, the condition should explicitly require either (a) demonstrating online detection within acceptable latency on new objects/grasps, or (b) reframing the contribution as 'offline-annotated tangible interaction' rather than object-agnostic interaction.","tokens_in":25252,"tokens_out":6690,"duration_ms":67509,"concrete_test":"Run an end-to-end trial with three new everyday objects not used in Studies 1-2. For each object, after a single egocentric view, execute the full detection pipeline under a 5-second wall-clock budget, with no offline scanning and no per-pose caching, then have participants perform the 0D/1D/2D tasks using self-chosen grasps. If task completion falls substantially below Table 1's 0% failure rate, or if detection latency exceeds the budget, the object-agnostic runtime claim is not supported by the current evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ATOM's central claim is object-agnostic tangible interaction, but the system as described is not object-agnostic at use time. The supplementary 'Setup requirements' state that each object must be scanned offline with RealityKit Object Capture (~10h, object static and visible), and that for each grasp pose, geometric element detection runs offline (~20s 2D refinement, ~2min detection, ~10s usability) before elements are cached; the online loop only performs contact detection on cached elements. User Study 1 (Section 6.1.1) instructed participants to grasp objects in predefined poses, so detection had already succeeded offline for those poses. Study 2 used only two fixed poses per object. Consequently, the 100% task completion in Table 1 and the SUS/TLX superiority demonstrate that interaction with pre-detected, cached elements works under controlled postures; they do not demonstrate that the detection pipeline generalizes to novel objects or arbitrary grasps at runtime. The Limitations section explicitly concedes that the geometric-element detection pipeline is 'not real time' and that tracking is susceptible to occlusion. The abstract's 'object-agnostic' is therefore an aspiration, and the strongest empirical claim is contingent on offline successes that the user studies never exercise. This is a scope mismatch between the central claim and the evidence, not an internal inconsistency; the ablation comparisons may still be valid in the cached-element setting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ATOM, a system that turns everyday handheld objects into AR tangible interfaces by detecting fine-grained geometric elements (corners, edges, and surfaces) in fingertip-reachable regions and mapping them to 0D, 1D, and 2D microgestures. The pipeline combines hand/object pose tracking, Canny edge detection on rendered normal maps, generative 3D enhancement via Hunyuan3D, generative 2D refinement via Gemini sketch simplification with ORB/RANSAC alignment, MANO-based reachability modeling, a five-factor usability scoring model, and a contact/voting interaction loop. The authors report two user studies: Study 1 (12 participants, within-subjects, 3 objects, 3 tasks, 4 trials per cell) compares the full pipeline against three ablations and reports 0% failed trials for the full system, significantly better completion times for the 1D and 2D tasks, and higher SUS/lower NASA-TLX scores; Study 2 extends to 10 objects under two poses and reports low failure rates; a system-level evaluation reports tracking stability and detection accuracy. The title and abstract claim 'object-agnostic' tangible interaction.","tokens_in":25528,"tokens_out":8453,"duration_ms":78657,"significance":"If supported, ATOM would be a useful contribution to opportunistic tangible AR interaction: the unified 0D/1D/2D metaphor over local geometry is elegant, and using generative foundation models to regularize noisy meshes and edge maps is a plausible way to avoid per-object training. The ablation study is well designed (within-subjects, counterbalanced, 1,728 planned trials), and the authors are admirably explicit about limitations, including non-real-time detection and occlusion sensitivity. However, the central 'object-agnostic' claim is materially narrower than what is demonstrated: because every object must be scanned offline (~10h) and per-pose detection runs offline (~2min) before elements are cached, the user studies exercise interaction with pre-computed caches under fixed poses, not runtime discovery on novel objects. The detection-accuracy evidence also lacks objective ground truth and inter-rater reliability metrics. These issues do not invalidate the cached-element framework, but they require either new evidence or a re-scoped claim.","major_comments":[{"comment":"The title and abstract claim 'object-agnostic' tangible interaction, but the system as evaluated is not object-agnostic at use time. The supplementary setup states that each object must be scanned offline with RealityKit Object Capture (~10h, object static and visible), and that for each grasp pose, 2D refinement (~20s), detection (~2min), and usability analysis (~10s) run before elements are cached; the online loop only performs contact detection on cached elements. Study 1 instructed participants to grasp objects in predefined poses, so detection had already succeeded offline for those poses, and Study 2 used only two fixed poses per object. The 100% task completion in Table 1 and the SUS/TLX results therefore demonstrate reliable interaction with pre-detected, cached elements under controlled postures, not generalization to novel objects or arbitrary grasps at runtime. The Limitations section concedes that detection is 'not real time.' The 'object-agnostic' claim should be re-scoped to pre-scanned objects after offline detection, or the paper should add an evaluation in which novel objects and unconstrained grasps are processed without pre-caching.","section":"Abstract; Supplementary §3 'Setup requirements'; §6.1.1"},{"comment":"The central detection-accuracy claim rests on an unvalidated subjective assessment. The paper states that 'two independent users assess whether each detected element in User Study 2 is correct' and reports 'around 62%' before and 'above 90%' after usability analysis, but it gives no definition of 'correct,' no per-object/per-pose counts, no confidence intervals, and no inter-rater reliability statistic such as Cohen's kappa. Without agreement information, the two raters may simply reproduce the authors' criteria, and 'above 90%' cannot be assessed. Because Q1 in Study 1 is specifically about detection accuracy and the objective 0D accuracy is not significantly different from the best baseline (p>0.05), the paper should supply an objective or grounded accuracy protocol, and should report detection rates per object and per DoF.","section":"§6.3 'Detection accuracy'; §6.2"},{"comment":"The usability weights and hard thresholds in the scoring function are load-bearing for the 'ours w/o UA' ablation, but their derivation is not reported in a way that prevents overfitting concerns. The supplementary states that the weights are 'empirically set ... from preliminary experiments' (w_rch=3.0, w_shp=3.0, etc.), and the thresholds (60% blocked, 30° sharpness, 1 cm, 4 cm²) appear hand-picked. If these values were tuned on the same objects and interaction tasks used in Study 1, the comparison between the full system and the no-UA baseline partly reflects the tuning rather than the general principle of usability analysis. The paper should report the preliminary experiment and tuning procedure, and ideally a sensitivity analysis or cross-validation across objects; at minimum it should state which objects and poses were used to set the weights and thresholds.","section":"Eq. (6); Supplementary 'Hyperparameters'"},{"comment":"The statistical reporting is incomplete in ways that affect the strength of the headline claims. Table 1 gives failure percentages without absolute counts or denominators, and no inferential test is applied to failure rates despite the large differences (e.g., 69.79% vs 0.00% in the 2D task). The completion-time analyses exclude skipped trials as timeouts, which can bias comparisons when failure rates differ across conditions. The paper also reports 'all with p<0.05' for the custom Likert questions in Figure 13 without giving test statistics, means, or SDs. In addition, the 1D comparison with 'ours w/o UA' is not significant (p=0.244) and the 0D accuracy is not significantly different from the best baseline, so the Discussion should state those outcomes more cautiously. A complete reporting of the RM-ANOVA results, effect sizes, and failure-rate tests, or an explicit labeling of failure rates as descriptive, is needed to support the 'full pipeline outperforms ablations' conclusion.","section":"§6.1.2; Table 1"}],"minor_comments":[{"comment":"The procedure says target values were randomized, but the supplementary lists fixed arrays of initial/target timestamps and pan targets; please clarify whether the values were randomized per participant or fixed with trial order randomized.","section":"§6.1.1; Supplementary §2"},{"comment":"Equation (1) is described as the proportion of rays blocked, but the formula sums ray hits without dividing by N×M; either normalize the expression or revise the description.","section":"Eq. (1)"},{"comment":"The label 'Warping Affinement' appears to be a typo for 'Warping Alignment' or 'Affine Refinement'; please correct.","section":"Figure 5"},{"comment":"The tracking robustness protocol measures pose stability with the object stationary against a moving headset; this does not directly quantify occlusion by the interacting hand, so the claim that tracking stays within contact thresholds at <50% occlusion should distinguish head-motion stability from hand-occlusion robustness.","section":"§6.3"},{"comment":"Study 2 reports only aggregate failed-trial rates and means; per-object results would help identify which objects and grasps are easy or hard and would strengthen the generalizability claim.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to generate strong interest at UIST because of the system integration and the bold 'object-agnostic' framing. My main concern is not the technical feasibility of the cached-element system but the gap between the title/abstract and the offline pre-scan/per-pose detection pipeline; if the authors can re-scope their claims or provide an online generalizability study, the paper would be much stronger. The detection-accuracy evaluation and hyperparameter reporting also need strengthening before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid systems paper with a real integration contribution, and the ablation results support the pipeline's components. The main caveat is that the 'object-agnostic' framing outruns the evidence: at use time every object is pre-scanned and per-pose detection is cached offline, so the user studies exercise cached elements, not runtime detection on novel objects. That is a scope mismatch, not a fatal flaw; the detection pipeline itself is evaluated separately, and the limitations section is honest about it.\n\nWhat's new: ATOM combines generative 3D enhancement, generative 2D edge refinement, fingertip reachability modeling, and a multi-factor usability ranking to find corners, edges, and surfaces for 0D, 1D, and 2D microgestures. Compared with Ubi Edge and AdapTUI, this is a genuinely new integration and a sensible extension of the fine-grained-element idea. The ablation design is the right way to validate it: the full pipeline's 0% failed trials in Study 1 and the significant SUS/TLX gains over the best ablation are credible, and the per-object consistency is a nice touch.\n\nSoft spots, in order. First, the 'object-agnostic' claim: the supplementary setup requirements state that each object is scanned offline and each grasp pose runs offline detection before caching, and the studies use predefined poses. So the headline claim is aspirational; the evidence is for offline detection plus cached interaction. The stress-test note is right about this. Second, detection accuracy in Study 2 rests on two independent assessors' judgment, with no inter-rater reliability metric and no ground truth; the 62% to 90+% numbers are only as strong as that protocol. Third, the usability weights and thresholds were set from preliminary experiments with no sensitivity analysis, so the comparison against 'w/o UA' could partly reflect tuning; this is mild, but a small sensitivity check would help. Fourth, the completion-time analysis for the 1D and 2D tasks excludes skipped trials while some baselines have very high skip rates (up to 69.79%), which can bias comparisons. Fifth, no code or data is released.\n\nStill, the central argument holds within the cached-element setting: each generative component helps in the ablations, and the usability ranking sharply reduces failures. The citation pattern is appropriate, and the paper engages with the relevant prior work. This is for HCI and AR researchers working on tangible interaction, on-object microgestures, and opportunistic controls. My recommendation: send it to peer review, with comments asking for a sensitivity analysis on weights, an IRR metric for detection assessment, a more careful treatment of skipped trials, and a toned-down 'object-agnostic' claim or an explicit statement of the system's current boundary.","headline":"A genuinely useful integration of foundation models and human factors for on-object microgestures; the ablation story is credible, but the 'object-agnostic' claim should be scoped to offline-cached elements.","tokens_in":26170,"tokens_out":2093,"would_cite":true,"duration_ms":22349,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Corners, edges, and surfaces on everyday objects can become tap, slide, and swipe controls without per-object authoring.","keywords":["tangible interaction","microgesture","augmented reality","geometric elements","generative 3D enhancement","generative 2D refinement","usability analysis","object-agnostic interaction"],"falsifier":"Scan a small object with one fine feature deliberately occluded, run the full pipeline, and measure the Chamfer distance between the pre- and post-enhancement geometries over that feature; if the displacement exceeds the system's own 0.5 cm fingertip-contact threshold, the detected tap or slide target no longer overlaps the physical one, and the reported task-completion benefit should disappear in a replication.","tokens_in":24929,"feed_emoji":"🖐️","tokens_out":10042,"duration_ms":94604,"temperature":0.7,"pith_summary":"ATOM claims that the small geometric features already present on everyday objects—protruding corners, sharp edge segments, and flat surface patches—are enough to turn any handheld object into a tangible interface, without matching the prop to a virtual counterpart. The paper builds a fingertip-aware detection pipeline that finds these fine features only where a finger can reach, cleans noisy object geometry with generative 3D and 2D models, and then ranks the detected elements by ergonomic usability. In the main ablation study the full pipeline completed all trials and beat each ablated version on task completion, system usability, and workload, and a second study kept low failure rates across ten objects. If the claim holds, AR users could control digital functions by tapping, sliding, and swiping the objects already in their hands, with no per-object authoring.","feed_headline":"Corners, edges, and surfaces turn any held object into an AR control","feed_subtitle":"ATOM maps tap, slide, and swipe gestures onto geometry already in your hand","key_machinery":"The carrying mechanism is the fingertip-aware detection-and-ranking pipeline. It restricts the search for interaction elements to a reachable volume computed from a parameterized hand model with collision filtering, which both shrinks the problem and makes the results physically usable. Noisy scans are cleaned twice: a generative 3D mesh model sharpens the geometry, and a generative image model simplifies the 2D edge map, with the simplified sketch warped back into alignment so detected elements stay on the physical object. A weighted usability score then discards elements that are blocked, ambiguous, too small, or too smooth, and selects the top corner, edge, and surface for each degree of freedom. This chain of reachability, enhancement, and ergonomic ranking is what turns raw geometry into controls a user can reliably operate.","core_discovery":"The paper sets out to show that fine-grained local geometry is a universal tangible affordance: protruding corners, sharp edge segments, and flat surface patches on an everyday object can be mapped to 0D taps, 1D slides, and 2D swipes, so no object-level match between a physical prop and a virtual controller is needed. ATOM detects those elements only inside the fingertip's reachable space, enhances the scanned object with a generative 3D mesh model, cleans the resulting edge map with a generative image model, and then ranks candidates by reaching cost, accessibility, hand and fingertip ergonomics, gesture ambiguity, and geometric sharpness. The empirical claim is that this full pipeline completed every trial in the main user study, outperformed all three ablations on subjective usability and workload, and produced low failure rates across ten objects in a second study.","pith_inferences":["A testable extension the paper does not run: extract the same corner–edge–surface vocabulary from the AR device's live scene mesh instead of an offline pre-scan; if detection accuracy stays near the reported level, the system would approach true object-agnostic interaction instead of object-set-specific interaction.","An inference left implicit: the five usability factors are independent of the generative stages, so they could rank elements from any future detector, including learned ones; the ergonomic formulation may be the most portable part of the contribution.","The paper's 0D results—comparable accuracy with a reachable but geometrically wrong corner—suggest that for discrete taps, usability filtering matters more than geometric correctness, which implies a redesign could trade detection precision for speed on 0D controls."],"forward_implications":["With the full pipeline, users completed 100 percent of trials across all three tasks in Study 1, while every ablation produced failures, so each generative stage and the usability filter earns its place.","Completion times around 4.91 seconds for 1D seeking and 7.53 seconds for 2D panning, with the smallest across-object variance of any condition, imply the same gesture vocabulary behaves consistently on objects of different shape and size.","SUS of 74.75 versus 54.50 and NASA-TLX of 27.50 versus 41.17 over the best baseline indicate the detected elements feel more usable and less effortful, not merely more detectable.","Failure rates of 6.67, 1.67, and 2.43 percent across ten objects in Study 2 support transfer across objects and grasps once each object has been pre-scanned."],"supporting_citations":[{"why":"Provides the generative 3D mesh model used to sharpen noisy object geometry in the enhancement stage.","marker":"[42]"},{"why":"Provides the generative image model prompted to simplify the edge map in the 2D refinement stage.","marker":"[17]"},{"why":"Supplies the prior edge-based tangible TUI approach that ATOM replaces manual line annotation with for automatic detection.","marker":"[22]"},{"why":"Supplies the environment-scale geometric element detection baseline that ATOM contrasts with on small everyday objects.","marker":"[21]"},{"why":"Supplies the parameterized hand model used to compute fingertip reachability and hand pose.","marker":"[54]"},{"why":"Supplies the twist-splay-bend frame used to penalize unnatural whole-hand poses in the ergonomics score.","marker":"[76]"},{"why":"Supplies the iterative closest point registration used to align the generated 3D mesh back to the scanned object.","marker":"[4]"},{"why":"Supplies ORB feature matching plus RANSAC to warp the refined edge map into spatial alignment with the original.","marker":"[55]"},{"why":"Supplies the corner detector used to locate corner candidates on the refined local edge map.","marker":"[20]"},{"why":"Supplies the density-based clustering used to consolidate corner candidates into stable tap targets.","marker":"[11]"}],"fun_headline_variants":["Fingertip-aware AR maps taps and swipes to any object's geometry","Corners become buttons: ATOM turns any object into a tangible UI","Local geometry becomes a 0D, 1D, or 2D gesture surface for AR","ATOM leverages corners, edges, and surfaces for object-agnostic interaction","Everyday objects' geometry drives AR microgestures without matching"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The object-agnostic claim rests on the assumption that each object can be scanned once in advance and that the two generative cleaning steps preserve, rather than distort, the small corners, edges, and surfaces the interactions depend on.","fun_headline_variants_meta":{"raw":{"variants":["Fingertip-aware AR maps taps and swipes to any object's geometry","Corners become buttons: ATOM turns any object into a tangible UI","Local geometry becomes a 0D, 1D, or 2D gesture surface for AR","ATOM leverages corners, edges, and surfaces for object-agnostic interaction","Everyday objects' geometry drives AR microgestures without matching"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1313,"prompt_tokens":895,"completion_tokens":418,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":511,"tokens_out":418,"duration_ms":4518,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:23:14.757018+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Scan a small object with one fine feature deliberately occluded, run the full pipeline, and measure the Chamfer distance between the pre- and post-enhancement geometries over that feature; if the displacement exceeds the system's own 0.5 cm fingertip-contact threshold, the detected tap or slide target no longer overlaps the physical one, and the reported task-completion benefit should disappear in a replication.","supporting_citations":[],"review_version":1}