{"id":"00954048-b471-4007-91eb-18e6afa3b5e9","arxiv_id":"2606.25162","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"fARfetch integrates shared landmark mapping, miniature path authoring, and VLM-based AR view management for collocated human-robot collaboration in large visually diverse environments, with significant gains shown in a real-world outdoor user study.","lead":"fARfetch is an AR-HRC system using shared semantic mapping, a world-in-miniature view, and VLM-driven adaptation of virtual content color, size, and orientation to support collaboration in large outdoor spaces. A 13-person study on a 30.5m inspection task showed 66% faster completion and lower workload than a non-AR baseline.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"VLM adaptation reliability across outdoor visual conditions is the least-secured assumption for the legibility and performance claims","rationale":"The reader's weakest_assumption directly identifies the same unverified link between VLM behavior and the reported user-study outcomes; no stronger internal inconsistency or missing control appears from the supplied abstract.","tokens_in":1747,"tokens_out":307,"duration_ms":13551,"concrete_test":"In the full methods and results sections, locate any quantitative VLM metrics (e.g., adaptation success rate, mean latency, or error counts per condition); recompute the completion-time and workload statistics after excluding trials where adaptation latency exceeded 500 ms or produced visible errors. If the significance or effect sizes change materially, the integrated-system claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline results (66% faster completion, workload reductions, and \"effectively maintained\" legibility) depend on component (iii) — VLM-driven joint adaptation of color/size/orientation — functioning without unacceptable latency or errors over the 30.5 m outdoor range. The abstract supplies only a custom survey outcome and does not report VLM accuracy, failure rate, or per-trial latency under the actual lighting/vegetation/background diversity encountered. If adaptation introduced delays or mis-adaptations on even a subset of trials, the observed time and workload gains could be attributable to the shared mapping and miniature components alone rather than the full fARfetch pipeline.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents fARfetch, an AR-HRC system for collocated collaboration in large visually diverse outdoor environments. It integrates (i) shared semantic environment mapping between AR headset and robot for landmark-grounded commands, (ii) a context-aware world-in-miniature for path authoring, and (iii) VLM-driven joint adaptation of virtual content color, size, and orientation to preserve legibility. Implemented on a Meta Quest 3 and Unitree Go2, a within-subjects user study (N=13) on a real 30.5 m outdoor inspection task reports 66% faster completion times versus a non-AR baseline, workload reductions (mental demand -43%, temporal demand -34%, frustration -66%), and positive results on a custom legibility survey.","tokens_in":1894,"tokens_out":614,"duration_ms":18152,"significance":"If the VLM adaptation component functions reliably, the work offers a practical contribution to outdoor AR-HRC by addressing legibility at distance and under visual variation. The real-hardware, outdoor evaluation with statistically significant time and workload gains provides ecological validity that is uncommon in AR robotics studies. The combination of mapping, miniature, and adaptive view management could inform systems for inspection, search-and-rescue, and field robotics where operators must maintain awareness beyond line-of-sight.","major_comments":[{"comment":"User Study Results: The central performance claims (66% faster completion and workload reductions) are presented as resulting from the full fARfetch pipeline, yet the manuscript reports no quantitative VLM metrics—adaptation accuracy, failure rate, or latency—under the actual outdoor lighting, vegetation, and background conditions of the 30.5 m task. This omission leaves open the possibility that observed gains derive primarily from components (i) and (ii) rather than the VLM adaptation in (iii).","section":"User Study Results"},{"comment":"VLM-Driven AR Content Adaptation section: The legibility claims rest on a custom survey outcome, but the paper provides no description of the VLM prompting strategy, model choice, or handling of edge cases (e.g., low light, high contrast vegetation). Without these details or failure-mode analysis, it is difficult to assess whether the adaptation introduces unacceptable latency or errors across the tested visual diversity.","section":"VLM-Driven AR Content Adaptation"}],"minor_comments":[{"comment":"Abstract and Results: The abstract states results are 'significantly' different but omits the statistical test, degrees of freedom, and exact p-values; these should be supplied for reproducibility.","section":"Abstract"},{"comment":"Implementation: The description of the shared mapping and miniature components would benefit from a brief diagram or pseudocode showing data flow between headset and robot to clarify how semantic landmarks are synchronized.","section":"Implementation"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. The comments highlight important areas for clarification regarding the VLM component's evaluation and implementation details. We address each major comment below and commit to revisions that strengthen the paper without misrepresenting the current work.","responses":[{"response":"We agree that the user study evaluates the integrated fARfetch system against a non-AR baseline and does not provide isolated quantitative metrics for the VLM adaptation component. The 66% time reduction and workload improvements are reported for the complete pipeline, which is consistent with the ecological validity goal of the outdoor evaluation. However, this leaves the specific contribution of component (iii) unquantified. In the revision, we will add a new subsection reporting VLM-specific metrics collected during the study (adaptation accuracy, failure rate, and latency) under the actual 30.5 m outdoor conditions to better attribute the observed gains.","revision_made":"yes","referee_comment":"[User Study Results] User Study Results: The central performance claims (66% faster completion and workload reductions) are presented as resulting from the full fARfetch pipeline, yet the manuscript reports no quantitative VLM metrics—adaptation accuracy, failure rate, or latency—under the actual outdoor lighting, vegetation, and background conditions of the 30.5 m task. This omission leaves open the possibility that observed gains derive primarily from components (i) and (ii) rather than the VLM adaptation in (iii)."},{"response":"We accept that the current manuscript omits key implementation details of the VLM-driven adaptation. The legibility survey results are presented without supporting technical description. In the revised manuscript, we will expand the VLM-Driven AR Content Adaptation section to specify the VLM model, the prompting strategy for jointly adapting color, size, and orientation, and include a failure-mode analysis drawn from the outdoor trials (including low-light and vegetation contrast cases) along with measured latency.","revision_made":"yes","referee_comment":"[VLM-Driven AR Content Adaptation] VLM-Driven AR Content Adaptation section: The legibility claims rest on a custom survey outcome, but the paper provides no description of the VLM prompting strategy, model choice, or handling of edge cases (e.g., low light, high contrast vegetation). Without these details or failure-mode analysis, it is difficult to assess whether the adaptation introduces unacceptable latency or errors across the tested visual diversity."}],"tokens_in":1528,"tokens_out":524,"duration_ms":24828,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main point is a working outdoor system that links a robot and AR headset through semantic maps, adds a miniature view for path planning, and uses a VLM to tweak AR object color, size, and orientation on the fly. The within-subjects study with 13 people on a 30.5 m inspection task reports 66% faster completion and clear drops in mental demand, temporal demand, and frustration versus a non-AR baseline.\n\nThe implementation on a Quest 3 and Unitree Go2 is concrete, and the task is realistic enough to matter for field work. The shared mapping and miniature components look like they address real coordination problems at distance.\n\nThe soft spot is the VLM adaptation. The abstract gives a custom legibility survey that came out positive, but nothing on how often the model picked the right adjustments, what the latency was, or how it handled the actual range of lighting and backgrounds in the trials. Without those numbers it is hard to tell whether the performance edge came from the adaptation or just from the mapping and miniature parts.\n\nN=13 is on the small side for claiming broad usability, and the custom survey leaves open questions about how the questions were worded and scored. The citation list in the abstract seems focused on prior AR-HRC work, which is fine for a systems paper.\n\nThis is for people building AR tools for outdoor robot teams. It is worth sending to peer review because the hardware setup and task are solid and the measured differences are large enough to be worth checking in detail. Reviewers will probably press on the VLM metrics and the survey design, but the core empirical result is there to discuss.","headline":"fARfetch puts together shared mapping, miniature authoring, and VLM adaptation for outdoor AR-HRC and shows time and workload gains in a 30m real-world study, but the VLM piece has almost no supporting numbers.","tokens_in":2376,"tokens_out":430,"would_cite":false,"duration_ms":8205,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"fARfetch uses vision-language models to adapt AR visuals so humans and robots can collaborate effectively across large outdoor spaces.","keywords":["augmented reality","human-robot collaboration","vision-language models","outdoor environments","AR content adaptation","shared semantic mapping","legibility","user study"],"falsifier":"A direct test showing whether legibility scores drop or error rates rise when the same 30.5 m task is repeated under extreme lighting shifts such as full sun versus deep shadow.","tokens_in":2660,"feed_emoji":"🥽","tokens_out":685,"duration_ms":16018,"temperature":0.7,"pith_summary":"The paper introduces fARfetch as an AR system for human-robot collaboration that adds shared semantic mapping of landmarks, a miniature world view for path planning, and automatic adjustment of virtual content via a vision-language model. In a study with 13 participants performing a 30.5-meter outdoor inspection task, the system produced 66 percent faster completion times and lower reported mental demand, temporal demand, and frustration compared with a non-AR baseline. The adaptation keeps overlaid information readable despite changing backgrounds and long distances. A sympathetic reader would care because outdoor settings have long blocked wider use of AR for directing robots in real work.","feed_headline":"AR system with VLM adaptation speeds outdoor robot tasks by 66%","feed_subtitle":"View management keeps virtual overlays readable across 30 m in changing outdoor conditions while cutting mental workload and frustration.","key_machinery":"VLM-driven AR view management that jointly adapts virtual content color, size, and orientation to maintain legibility.","core_discovery":"The paper establishes that a combination of shared semantic environment mapping, a context-aware world-in-miniature interface, and vision-language-model-driven adaptation of AR content color, size, and orientation enables collocated human-robot collaboration to remain usable in large visually diverse outdoor environments, as shown by significantly improved task speed and reduced workload in a real-world 30.5 m inspection study.","pith_inferences":["The same adaptation loop could be applied to indoor scenes with rapidly changing lighting or to mobile robots operating in construction zones.","Removing the need for manual AR tuning might let non-expert users direct robots in new environments without prior calibration.","Extending the shared mapping to include dynamic objects could support collaboration in settings where both people and robots move continuously.","If the VLM adaptation proves robust, similar view-management logic might transfer to other mixed-reality interfaces that must handle scale and visual diversity."],"forward_implications":["Landmark-grounded go-to commands become usable because detected landmarks appear as AR anchors visible to both human and robot.","Fine-grained path authoring is supported through the miniature representation without requiring the operator to walk the full route.","Virtual overlays stay readable at long distances and across varied backgrounds, removing a key barrier to outdoor AR-HRC.","Overall operator workload decreases measurably in mental demand, temporal demand, and frustration during extended tasks."],"fun_headline_variants":["fARfetch speeds outdoor AR robot tasks 66% with VLM","VLM adapts AR color size for legible robot views at 30m","Shared semantic maps enable AR robot commands in outdoors","AR system cuts mental demand 43% in large inspection task","World miniature lowers frustration 66% for collocated HRC"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The vision-language model can adapt AR content to preserve legibility without introducing unacceptable latency or errors across the range of outdoor visual conditions.","fun_headline_variants_meta":{"raw":{"variants":["fARfetch speeds outdoor AR robot tasks 66% with VLM","VLM adapts AR color size for legible robot views at 30m","Shared semantic maps enable AR robot commands in outdoors","AR system cuts mental demand 43% in large inspection task","World miniature lowers frustration 66% for collocated HRC"]},"model":"grok-4.3","cost_usd":0.005251,"raw_usage":{"total_tokens":2552,"prompt_tokens":688,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":52512000,"prompt_tokens_details":{"text_tokens":688,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1779,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":688,"tokens_out":85,"duration_ms":10521,"temperature":1.0,"reasoning_tokens":1779,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-25T23:43:02.346337+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A direct test showing whether legibility scores drop or error rates rise when the same 30.5 m task is repeated under extreme lighting shifts such as full sun versus deep shadow.","supporting_citations":[],"review_version":1}