{"id":"94ed6cb6-72af-4f10-a819-e88a43d4fa0f","arxiv_id":"2608.12759","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"A seven-day diary study of 12 low-vision participants found that a customizable mobile AR app that highlights curbs, sidewalks, vehicles, and other outdoor objects can aid real-world navigation, but sunlight, rain, shadows, and nonstandard markings degrade recognition and augmentation visibility.","lead":"This paper reports a seven-day diary study in which 12 people with low vision used NavSight, a smartphone AR app that outlines and highlights outdoor objects to support navigation. It documents how the app helped users in daily life and where it failed, including glare, rain, shadows, and unusual road markings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Environmental degradation of recognition rests on self-attribution, not logged model outputs; controlled replication would settle it.","rationale":"The reader's weakest assumption (self-report bias) is correct, but the most load-bearing edge is narrower: the paper's recognition-degradation findings under weather, lighting, and nonstandard textures are presented as system-level findings even though no recognition outputs were logged. This is explicitly conceded in Section 6.4. The qualitative findings on usage patterns, configuration strategies, error interpretation, and social acceptability are well supported by the diary data and do not depend on objective recognition logs, so the core contribution remains intact. A controlled validation of the same model under the named conditions would settle whether the environmental-factor claim holds; if it fails, Section 6.3.2 and the abstract need wording calibration, but the central empirical account of user experience would not collapse. Because the authors already attribute these observations to participants in Section 5.4.2 and transparently disclose the logging limitation, the ACCEPT verdict stands; I would only request a small revision to align the abstract's causal phrasing with the evidence base.","tokens_in":34220,"tokens_out":8099,"duration_ms":93031,"concrete_test":"Run a short controlled outdoor evaluation with the same YOLO11l-seg-outdoor model on matched routes containing (a) dry sidewalk, (b) wet sidewalk or puddles, (c) tree-shadowed sidewalk, and (d) nonstandard crosswalk markings, with ground-truth labels. Compute per-condition FNR and FDR for sidewalk, curb, and crosswalk, and test whether the differences are significant. If condition-specific accuracy does not differ, the environmental-factor claim should be reframed as user-perceived; if it does, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 6.3.2 state that weather, lighting/shadows, and nonstandard markings degrade NavSight's recognition. The supporting evidence is participants' retrospective attributions plus a few illustrative screenshots (Section 5.4.2, Figure 8), not logged recognition outputs. Section 6.4 explicitly concedes that camera feeds and recognition results were not logged to preserve privacy. The causal link between environmental conditions and model failures is therefore not established: users may attribute errors to salient conditions such as rain or shadows even when the actual cause is viewpoint, distance, or object ambiguity. This matters because the design implications in Section 6.3.2 (weather-robust datasets, image enhancement before inference) and the strongest claim that external conditions 'degrade recognition' depend on that link. The safety line 'no participant reported a fall or injury' is also weak given 12 participants and roughly one week, but it is secondary and the paper does not lean on it heavily.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents NavSight, a mobile AR application for people with low vision that recognizes 21 outdoor object categories and renders six customizable visual augmentations in real time on an iPhone. The recognition component is a YOLO11l-seg model fine-tuned on Mapillary Vistas and evaluated on a static test split (mAP50=0.588, mAP75=0.305, mAP=0.318, with per-class AP/FNR/FDR reported in Appendix A). The authors deployed NavSight in a seven-day diary study with 12 participants with low vision, collecting daily 5-point ratings, privacy-preserving usage logs, user-captured screenshots, and semi-structured exit interviews. The findings characterize how participants used the app across outdoor scenarios, how augmentations changed scene perception (walkable versus non-walkable regions, tripping hazards, moving objects, visual reach), how the phone competed for attention and hands, how participants interpreted and coped with recognition errors, how they configured object selection and augmentation over time, and how they experienced social acceptability and usability in the wild. The paper derives design implications for context-aware configuration, debugging support, platform options, expanding recognition to surface-level hazards, supporting trust calibration, and consequence-aware evaluation metrics. Limitations, including self-reported measures and the absence of logged camera or recognition outputs, are acknowledged in Section 6.4.","tokens_in":34207,"tokens_out":11692,"duration_ms":112004,"significance":"If the findings hold, this is a valuable and timely contribution: it is one of the first in-the-wild studies of mobile AR visual augmentations for low vision outdoor navigation, and it provides ecological-validity evidence that short lab-based evaluations cannot supply. The paper is methodologically careful in several respects: the recognition model is quantitatively evaluated on a held-out test set with per-class metrics; the study triangulates daily ratings, usage logs, screenshots, and interviews; the logging design preserves bystander privacy; and the authors plan to open-source NavSight. The qualitative findings on configuration strategies, mental models of AI errors, and social acceptability are credible and generate concrete design hypotheses for future assistive AR systems. The main caveat is that the causal claims about environmental degradation of recognition rest on participant self-report rather than logged model outputs; this does not undermine the experiential findings, but the paper should label them as perceived effects and call for quantitative verification.","major_comments":[{"comment":"The paper presents as a finding that weather conditions, lighting and shadows, and nonstandard road markings degrade NavSight's recognition, and Section 6.3.2 builds model-training implications (weather-robust datasets, image enhancement before inference) on this claim. The supporting evidence, however, consists of participants' retrospective attributions and a few illustrative screenshots; Section 6.4 explicitly states that camera feeds and recognition results were not logged to preserve privacy. Without logged recognition outputs or a controlled replication, the causal link between environmental conditions and model failures is not established, because users may attribute errors to salient conditions such as rain or shadows even when the true cause is viewpoint, distance, or object ambiguity. I recommend consistently framing these as user-perceived degradations, including in the abstract, and explicitly stating that the causal mechanism requires quantitative verification before model-training implications are drawn.","section":"5.4.2 and 6.3.2 (also abstract)"}],"minor_comments":[{"comment":"Please reconcile the day counts: P9 is listed as using NavSight for 11 days, but Section 5.6.1 says she rated comfort 5 on eight of her nine days; P7 is listed as 8 days, but Section 5.6.2 says she rated discomfort on all seven days.","section":"Table 2 and Section 5.6"},{"comment":"The sentence 'no participant reported a fall or injury while using NavSight' should be qualified by the small sample size and short deployment period, since it is not strong evidence of safety on its own.","section":"5.2"},{"comment":"The deployment thresholds (confidence = 0.4, IoU = 0.5) appear only in a table note; please state and briefly justify these choices in the main text, since FNR and FDR depend on them.","section":"Section 3.4 and Table 4 note"},{"comment":"The claim that no research has investigated real-world feasibility and challenges of AR systems for low vision is a strong universal statement; consider softening it to 'to our knowledge, no prior work has...' to avoid overclaiming.","section":"Abstract and Section 1"},{"comment":"The claim that all participants converged on a stable set of augmented objects would be easier to assess with a small quantitative summary, such as the number of object-selection changes per participant per day, since the current support is narrative and spread across examples.","section":"5.5.1"}],"recommendation":"major_revision","confidential_remarks":"This is a strong empirical paper with a clear contribution. The major revision is about aligning the language of environmental effects with the self-report evidence; it does not require new data collection. I would encourage the authors to make the framing fix, reconcile the day-count inconsistencies, and resubmit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first field deployment I know of real-time visual AR augmentations for low-vision outdoor navigation, and it is a solid piece of empirical HCI. The study tracks 12 low-vision participants for a week with usage logs, daily ratings, screenshots, and exit interviews. The headline finding is not that the augmentations work, it's how people actually use them: they simplify scenes into walkable/non-walkable regions, notice tripping hazards, scan for traffic, and they keep reconfiguring object selection and augmentation styles as lighting and context change. That behavioral material is new and credible.\n\nThe paper earns credit for transparency. Section 6.4 concedes the small sample, iOS-only deployment, self-report measures, and the absence of camera/recognition logs. The recognition model is evaluated on a fixed test split with per-class AP/FNR/FDR, and the weak classes (sidewalk, curb, crosswalk) line up with what participants complained about. The 'holes in the augmentation' finding, where users detect unrecognized hazards as gaps in the sidewalk overlay, is a genuinely nice observation. The discussion of trust calibration and consequence-based error metrics is thoughtful rather than speculative.\n\nNow the soft spots, in proportion. The stress-test note is right: the paper states that weather, lighting/shadows, and nonstandard markings degrade recognition, but the evidence is participants' retrospective attributions plus a few screenshots, not logged model outputs. Section 6.4 explicitly says camera feeds and recognition results were not logged. So the causal link between environment and model failure is not established. Users could be misattributing errors to salient conditions. The abstract and Section 6.3.2 push the claim a bit harder than the evidence supports. That should be fixed either by logging or by reframing as perceived environmental challenges. It is a moderate limitation, not a fatal one. The central qualitative findings about usage and adaptation rest on the diaries and logs and hold up.\n\nMinor items: no inter-rater reliability for the thematic coding; the 'no falls/injuries' safety line is weak at this sample size and duration, but the paper does not lean on it heavily. Code and data are promised but not yet public, so reproducibility is partial.\n\nBottom line: this deserves a serious referee and is likely a solid venue paper. I would send it to review with a request to soften or substantiate the environmental-degradation causal claim and add coding reliability. I would cite it as the main real-world AR low-vision deployment.","headline":"A solid week-long field study of AR low-vision navigation; the environmental-degradation finding is real but self-reported and should be framed as user-perceived, not model-verified.","tokens_in":34886,"tokens_out":2143,"would_cite":true,"duration_ms":22097,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports a seven-day diary study in which 12 people with low vision used NavSight, a mobile AR app that augments recognized outdoor objects, and finds that the augmentations support navigation by simplifying scenes, while…","keywords":["low vision","augmented reality","outdoor navigation","diary study","visual augmentations","object recognition","accessibility","field deployment"],"falsifier":"Log the phone's camera feed, recognition outputs, and phone motion during the same diary protocol and compare them against participants' daily ratings and reported falls; if objective recognition failures and gait disruptions do not correlate with self-reported safety and helpfulness, or if falls are observed during sessions rated safe, the central characterization would need revision.","tokens_in":33884,"feed_emoji":"🧭","tokens_out":4659,"duration_ms":42697,"temperature":0.7,"pith_summary":"The paper seeks to establish that a mobile augmented-reality application can help people with low vision navigate uncontrolled outdoor environments, and to characterize how real-world use differs from lab evaluations. It presents NavSight, an app that recognizes 21 categories of outdoor objects and renders user-configurable visual augmentations, and reports a seven-day diary study with 12 low-vision participants. The study finds that augmentations simplified scenes into walkable and non-walkable regions, made tripping hazards and moving objects easier to notice, and extended visual reach, while also competing for attention and hands. Participants adapted their object selection, grouping, and augmentation styles across scenarios, and treated recognition errors differently depending on whether the error hid a hazard or merely added information. External conditions such as sun glare, wet surfaces, shadows, and nonstandard road markings degraded both recognition and augmentation visibility.","feed_headline":"Field test: AR app aids low-vision outdoor navigation","feed_subtitle":"Twelve low-vision users tested augmented walkways, curbs, and cars in daily life; weather and shadows broke recognition.","key_machinery":"The central object is NavSight, a smartphone AR app that runs a fine-tuned instance-segmentation model on the live camera feed and renders six user-configurable visual augmentations (contour enhancement, solid overlay, flashing, brightness adjustment, background darkening, and color removal) over two user-assigned object groups drawn from 21 outdoor object categories. The diary-study design—daily surveys, usage logs, screenshots, and a semi-structured exit interview—is the mechanism that moves the evaluation out of the lab and onto real sidewalks, driveways, and crossings. The load-bearing idea is that segmentation-based augmentations let users see environment structure directly, so even unrecognized hazards appear as holes in the augmented walkable surface.","core_discovery":"NavSight's real-time visual augmentations can support real-world outdoor navigation for people with low vision by simplifying scenes into walkable and non-walkable regions, making tripping hazards and moving objects easier to notice, and extending visual reach, but external conditions such as sun glare, wet surfaces, shadows, and nonstandard markings degrade recognition and augmentation visibility, and users adapt by reconfiguring object selection, grouping, and effects over time. The paper derives this from a seven-day diary study in which 12 low-vision participants used the app in their own neighborhoods, parking lots, crossings, parks, and even indoor malls, with daily surveys, usage logs, screenshots, and exit interviews as evidence. A related finding is that segmentation-based augmentations can reveal unrecognized hazards as gaps in the augmented walkable surface, giving users a way to notice objects the model cannot name.","pith_inferences":["A wearable or gaze-aligned display would likely reduce the reported attention competition, but the paper's own social-acceptability findings suggest it could increase conspicuousness and the risk of being seen as filming; this trade-off is testable by running the same diary protocol across form factors.","The reliance on self-report could be tightened by logging camera frames, recognition confidence, and phone motion, then checking whether daily helpfulness and safety ratings track objective recognition failures and gait disruptions.","The asymmetry in error tolerance implies that evaluation metrics for assistive AR should weight false negatives on hazards more heavily than mAP-style averages do; a cost-sensitive metric derived from navigation outcomes would be a concrete next step.","The finding that false positives on out-of-list objects were sometimes helpful suggests a design where uncertain detections are deliberately rendered in a distinct 'possible hazard' style rather than suppressed."],"forward_implications":["Designers of low-vision AR navigation aids should expect users to configure augmentations dynamically rather than use fixed settings, and should support those changes with low overhead.","Recognition quality must be judged by consequence: errors that hide hazards (e.g., a step marked as sidewalk) are dangerous, while false positives that merely add information are tolerated and sometimes useful.","Environmental conditions, especially wet surfaces, shadows, and nonstandard markings, are first-order failure modes for AI recognition and need explicit mitigation before such aids can be relied on outdoors.","A phone-based form factor trades social acceptability and availability for divided visual attention and occupied hands; users adapt but still report distraction, especially while crossing streets.","Segmentation displays can multiply their value by exposing unrecognized hazards as gaps in the augmentation, suggesting that surface-level hazard detection should be a priority."],"supporting_citations":[{"why":"Provides the contour-outline obstacle cueing design that NavSight adapts.","marker":"[34]"},{"why":"Contributes the Wizard-of-Oz AR obstacle highlighting approach that motivates NavSight's obstacle augmentations.","marker":"[74]"},{"why":"Supplies wayfinding guidance designs and evaluation practice for low-vision AR navigation.","marker":"[131]"},{"why":"Source for brightness and contour enhancement effects used in NavSight.","marker":"[132]"},{"why":"Source for flashing and contour cue designs that NavSight's augmentations build on.","marker":"[133]"},{"why":"Seven-day diary study method for deploying assistive technology with blind and low-vision users in the wild.","marker":"[136]"},{"why":"Prior field deployment of an AI-powered scene description app, informing the diary study approach.","marker":"[37]"},{"why":"Large open street-scene dataset used to fine-tune and evaluate NavSight's recognition model.","marker":"[82]"},{"why":"Base YOLO11 model that was fine-tuned into the outdoor object recognition model.","marker":"[54]"},{"why":"Documents daily outdoor navigation challenges of people with low vision, motivating NavSight's object categories.","marker":"[106]"}],"fun_headline_variants":["AR navigation app tested in the wild for low vision","Real-world test: AR app for low-vision navigation faces glare and shadows","Low-vision users adapt as AR navigation app breaks in sun and shadow","Seven-day diary: AR app for low vision survives sun, glare, and shadows","Real-world AR navigation for low vision: what worked, what broke"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on participants' self-reported daily ratings and retrospective interview accounts being accurate measures of helpfulness, safety, accuracy, and distraction.","fun_headline_variants_meta":{"raw":{"variants":["AR navigation app tested in the wild for low vision","Real-world test: AR app for low-vision navigation faces glare and shadows","Low-vision users adapt as AR navigation app breaks in sun and shadow","Seven-day diary: AR app for low vision survives sun, glare, and shadows","Real-world AR navigation for low vision: what worked, what broke"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000953,"raw_usage":{"total_tokens":4051,"prompt_tokens":921,"completion_tokens":3130,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":3036}},"tokens_in":537,"tokens_out":3130,"duration_ms":19607,"temperature":1.0,"reasoning_tokens":3036,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:53:46.443948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Log the phone's camera feed, recognition outputs, and phone motion during the same diary protocol and compare them against participants' daily ratings and reported falls; if objective recognition failures and gait disruptions do not correlate with self-reported safety and helpfulness, or if falls are observed during sessions rated safe, the central characterization would need revision.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the Wizard-of-Oz AR obstacle highlighting approach that motivates NavSight's obstacle augmentations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source for brightness and contour enhancement effects used in NavSight."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large open street-scene dataset used to fine-tune and evaluate NavSight's recognition model."}],"review_version":1}