{"id":"22a02dc1-69f6-4283-b0b8-d9dd752413f4","arxiv_id":"2506.09236","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review of 90 papers proposes a six-facet taxonomy and identifies 12 interface elements for AR user interfaces used by first responders.","lead":"This paper surveys 90 studies on augmented reality interfaces for first responders, covering EMS, firefighting, and law enforcement, and builds a six-facet classification of interface designs. It is a useful map for researchers and developers of what has been tried, what shows promise, and what remains untested in this safety-critical niche.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comprehensiveness claim is not yet supported: the 90-paper corpus comes from a four-database keyword search whose population terms omit 'search and rescue' and 'disaster response', with no citation chasing or saturation check, so the taxonomy and gaps may partly reflect search coverage.","rationale":"Reader's verdict is CONDITIONAL, and the weakest assumption is search coverage; my stress test agrees with that assessment rather than introducing a different objection. I identified the same load-bearing point because the 'comprehensive map' claim is the paper's headline novelty, and it rests entirely on the 90-paper sample. The manuscript is internally consistent about its method: PRISMA flow, keyword tables, and inclusion criteria are present, and the authors honestly disclose the search limitation. But the disclosure does not by itself settle whether the limitation is material. The concrete test above would settle it by checking whether a broader retrieval changes the taxonomy or the reported gaps. If the test shows material additions, the paper should be revised to claim a map of a keyword-searchable subset rather than a comprehensive map, and the gap analysis (e.g., sparse haptic/auditory feedback) should be re-validated. Because this is testable and the current evidence is insufficient to confirm or refute it, conditional acceptance remains the right call. A secondary observation is that the conclusion's list of '12 distinct interface elements' actually enumerates 11 items; this strengthens the case for releasing the coding scheme, but it is not the primary threat to the central claim.","tokens_in":28827,"tokens_out":11216,"duration_ms":97902,"concrete_test":"Re-run the search in the same four databases with the Table 1 string augmented by population terms 'search and rescue', 'rescue', 'disaster response', 'emergency management', 'paramedic', 'police', and 'emergency service', and add backward/forward citation chasing from all 90 included papers. Screen the new records with the stated inclusion criteria, then classify the additional papers into the Figure 4 facets and the Section 3.3 interface-element list. If any new paper requires a facet value or interface element not already in the taxonomy, or if the Table 7 haptic/auditory counts (2 and 10) shift by more than a few papers, the 'first comprehensive map' and the derived gap claims need explicit qualification. If the expanded corpus adds no new categories and leaves counts essentially unchanged, the concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 5: 'To our knowledge, this is the first scoping review to comprehensively map the domain of AR UIs for first responders') depends on the 90-paper corpus being representative of the field. That condition is weakly secured. Section 2.6 concedes that 'potentially relevant articles that did not match the search string were not included,' and the search terms in Table 1 are narrower than the domain: the Population block covers EMS/EMT, firefighting, law enforcement, and 'first responder/emergency responder/public safety' variants, but does not include 'search and rescue,' 'rescue,' 'disaster response,' 'emergency management,' 'paramedic,' or 'police' as standalone terms. A paper titled, say, 'Augmented Reality for Urban Search and Rescue' would not be retrieved unless it happened to mention one of the listed population terms in its title, abstract, or keywords. The protocol also reports no backward/forward citation chasing and no comparison with prior reviews such as [2], even though [2] is cited as a systematic review of situational-awareness technologies for disaster response. Consequently, the gap analysis (e.g., only 2 haptic and 10 auditory papers in Table 7; few cross-discipline designs in Section 4.2) could be an artifact of the retrieval strategy rather than a true property of the literature. This is a standard scoping-review limitation, but it bears directly on the 'comprehensive map' claim, which is the paper's main novelty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a scoping literature review of augmented-reality user interfaces (AR UIs) for first responders in the EMS, firefighting, and law-enforcement domains. Following Kitchenham and Charters and the PICOC framework, the authors searched ACM Digital Library, IEEE Xplore, ProQuest, and Scopus, screened candidate records, and retained 90 papers. They propose a six-facet taxonomy (operating environment, discipline, requirements, display hardware, display context, output channel) and a catalog of 12 interface elements (e.g., edge detection, object highlighting, navigation, alerts, on-demand interfaces, augmented guidance). The review reports publication trends, classifies all 90 papers in facet tables, and identifies gaps such as sparse multimodal feedback and limited cross-discipline designs.","tokens_in":29141,"tokens_out":6965,"duration_ms":66603,"significance":"If the corpus is representative, the review is a useful synthesis and appears to be the first dedicated map of this domain. Its strengths are the explicit PICOC/PRISMA procedure, dual screening and data-extraction, complete citation tables for every facet, and candid statements of limitations. The taxonomy is simple and usable, and the 12-element catalog gives designers a shared vocabulary. However, the significance depends on the representativeness of the 90-paper corpus and on the trustworthiness of the classifications; both are currently in question.","major_comments":[{"comment":"The central claim that this is the first comprehensive map of AR UIs for first responders is not yet supported by the retrieval strategy. The Population block of Table 1 omits high-frequency domain terms such as 'search and rescue', 'rescue', 'disaster response', 'emergency management', 'paramedic', and standalone 'police'; a paper titled, say, 'Augmented Reality for Urban Search and Rescue' would be retrieved only if it happened to include one of the listed population phrases. Section 2.6 concedes that non-matching articles were excluded, but no backward/forward citation chasing, no saturation check, and no comparison with the prior systematic review [2] is reported. Because the gap analysis (e.g., only 2 haptic and 10 auditory papers in Table 7; few cross-discipline designs in Section 4.2) is derived from this corpus, those gaps may partly reflect search coverage. The authors should either broaden the search (adding terms, reference harvesting, and a comparison against [2]) or temper the comprehensiveness claim to one about the retrieved corpus.","section":"§2.3, Table 1; §2.6; §5"},{"comment":"The PRISMA flow and the reported totals are inconsistent. The abstract and Section 2.4 say the keyword search retrieved 1,751 papers; Section 2.4 then says 'Of the initial 1,751 retrieved papers, 1,661 were rejected', which matches the eligibility stage in Figure 1. Figure 1, however, reports 3,440 records screened (after 1,169 duplicates removed) and 1,751 citations sought for retrieval. The reader cannot tell whether 1,751 is the raw retrieval, the post-deduplication count, or the number of full-text reports assessed. Please redraw the flow with consistent stage counts and define each number.","section":"§2.4, Fig. 1; Abstract"},{"comment":"The classification tables are not internally reproducible. Within a single facet, categories are not mutually exclusive: Display Hardware sums to 108 papers, Requirements to 99, Display Context to 118, and Output Modality to 102 against N=90, so a paper can appear in several categories. The text does not state that categories are non-exclusive, nor does it report how overlaps are handled or how disagreements between the two readers were resolved beyond discussion (§2.4–2.5). No inter-rater reliability statistic (e.g., Cohen's kappa) is reported for the classification. Because the taxonomy is the paper's main contribution, the authors should state the exclusivity rules, report agreement coefficients, and provide per-paper facet values in the supplementary material.","section":"§3.2, Tables 2–7; Fig. 4"},{"comment":"The search strategy is not fully replicable. The paper gives the synonym lists in Table 1 and states that the base string 'was adapted for the databases', but it does not report the actual query strings used for ACM, IEEE Xplore, ProQuest, or Scopus, nor the date of the search for each source. Full query strings (with field tags, wildcards, and filters) should be supplied in an appendix or online supplement.","section":"§2.3"}],"minor_comments":[{"comment":"The running header misspells 'Transactions' as 'TRANSACTIOINS'.","section":"Header"},{"comment":"Figure 4 uses 'Personal Records Data' while the text and Table 4 use 'Patient/Suspect Data'; the terms should be aligned.","section":"Fig. 4; §3.2.3; Table 4"},{"comment":"Figure 4's Output Channel branch lists 'Non-Visual' in addition to 'Visual', but Section 3.2.6 defines only Visual, Auditory, and Haptic and says all included interfaces have visual augmentation; clarify the figure.","section":"Fig. 4; §3.2.6; Table 7"},{"comment":"Table 5 labels a category 'HMD (Unspecified)' while Figure 4 says 'Head Mounted'; similarly, Table 6 says 'Spatial' while Figure 4 says 'Spatial Augmentation'. Use consistent facet names.","section":"Table 5; Fig. 4"},{"comment":"The statement that navigation elements were 'mentioned (103 instances across various disciplines)' uses an undefined counting unit and exceeds the 90-paper corpus; explain how instances are counted.","section":"§3.3.3"},{"comment":"Figure 2 has no visible y-axis label; 'Number of Publications' should appear in the figure itself.","section":"Fig. 2"},{"comment":"The sentence beginning 'Inclusion criteria we relevance' contains a typo; it should read 'Inclusion criteria were relevance, availability, and language (English)'.","section":"§2.4"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the manuscript's fit with TVCG's usual archival audience rests on the taxonomy's utility as a lasting contribution, since the paper contains no new technical results. A brief statement distinguishing the scoping-review contribution from the prior design work of the authors would be helpful: papers [37,38] are part of the reviewed corpus and their design concepts (on-demand interfaces, augmented enhancement) appear as taxonomy categories, so a short disclosure of that relationship is appropriate. No concerns about research misconduct are raised."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before you cite it as the canonical map of AR interfaces for first responders. The six-facet taxonomy and the 12-element interface catalog are genuinely new syntheses, and the classification tables (Tables 2-8) let you see exactly which paper lands where. That transparency is a real strength: the authors report their PRISMA flow, their base search string, and they list every paper per facet and per interface element. They also acknowledge in Section 2.6 that keyword searches miss relevant work. That limitation section is honest, not pro forma.\n\nThe main soft spot is the gap between the 'comprehensive map' claim and the actual retrieval strategy. The population terms in Table 1 omit 'search and rescue,' 'disaster response,' 'paramedic,' and standalone 'police,' and there is no backward/forward citation chasing and no comparison with the earlier systematic review they cite as [2]. A paper titled 'AR for Urban Search and Rescue' could easily slip through unless it happened to mention one of the listed population terms. So the gap analysis (e.g., only two haptic and ten auditory papers) may partly reflect search coverage rather than the literature itself. This is a standard scoping-review limitation, but it bears directly on the paper's central claim, so it deserves a serious fix: report the database-specific search strings, add citation chasing or at least compare against prior reviews, and soften the 'comprehensive' wording if the search is not expanded.\n\nTwo smaller points. First, the reader's worry about 'search strings and coded data not shipped' is only partly right: the base search string is in Table 1 and the per-paper classifications are in the tables, but the database-specific adaptations and inter-rater reliability statistics are missing. Second, Figure 4 lists 'Non-Visual' as an output channel, but the text and Table 7 only define visual, auditory, and haptic; that inconsistency should be cleaned up.\n\nThis is a solid, useful piece of work, not a revolutionary one. The taxonomy is a practical organizing device, and the gap analysis, once secured against search artifacts, will guide research scoping and design choices. It deserves a serious referee, and the right outcome is conditional acceptance with a request for the search protocol details and a modest toning down of the 'comprehensive' claim.\n\nRecommendation: send to peer review. Ask for the per-database search strings, inter-rater reliability or a justification for consensus coding, and either citation chasing or a revised claim.","headline":"A useful, competently reported scoping review whose new taxonomy and interface catalog are worth having, despite search-coverage gaps that soften the 'comprehensive map' claim.","tokens_in":29642,"tokens_out":2421,"would_cite":true,"duration_ms":26531,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ninety peer-reviewed papers on augmented-reality interfaces for first responders are organized into a six-facet taxonomy and a catalog of 12 recurring interface elements, giving the field a common design vocabulary and a list of gaps.","keywords":["augmented reality","first responders","scoping review","taxonomy","interface elements","situational awareness","public safety","head-mounted display"],"falsifier":"A replication that broadens the search—for example, adding Web of Science and PubMed, adding terms like 'wearable display' and 'see-through display,' and hand-searching the reference lists of the 90 included papers—would test the map's completeness. If such a search surfaced a sizeable set of relevant AR UI papers that do not fit into the six facets or 12 interface elements, or that substantially change the reported counts (e.g., many more haptic or auditory interfaces), the claim that Figure 4 maps the domain would be weakened.","tokens_in":28667,"feed_emoji":"🚨","tokens_out":3786,"duration_ms":34255,"temperature":0.7,"pith_summary":"The paper claims to be the first scoping review that comprehensively maps augmented-reality user interfaces for first responders. It screened 1,751 records from ACM, IEEE, ProQuest, and Scopus, kept 90 peer-reviewed papers, and organized their designs into a six-facet taxonomy: operating environment, discipline, data requirements, display hardware, display context, and output channel. Across those papers it also catalogued 12 recurring interface elements, from edge detection and X-ray views to navigation aids and on-demand menus. The review's value is that it gives researchers and developers a common vocabulary and a gap list: little multimodal feedback, sparse cross-discipline designs, and few rigorous evaluations of effectiveness.","feed_headline":"90 studies map AR interfaces for first responders into six facets","feed_subtitle":"Taxonomy spans environment, discipline, data, hardware, display, output; biggest gap is multimodal feedback.","key_machinery":"The carrying device is the faceted taxonomy, which classifies each system along six independent dimensions so that any combination of facet values describes a potential AR configuration; the paper pairs it with a catalog of 12 interface elements derived from the included studies. The taxonomy does the argument's work because it converts a scattered set of prototypes into a structured design space, and the gap analysis is read off that structure.","core_discovery":"The central claim is that the design space of AR user interfaces for first responders can be organized by six orthogonal facets—operating environment (field, command center, training), public-safety discipline (EMS, firefighting, law enforcement, other), primary data requirement (environment, physiological, patient/suspect, context-unaware), display hardware (head-mounted, handheld, stationary), display context (spatial augmentation vs heads-up display), and output channel (visual, with occasional auditory or haptic)—and that 12 interface elements (edge detection, object highlighting, X-ray, spatial reconstruction, gaze indicators, navigation aids, alerts, on-demand interfaces, annotations, augmented support, augmented enhancement, and vitals monitors) recur across the 90 papers. The paper argues that this taxonomy and element catalog constitute the first comprehensive map of the domain, that navigation interfaces—especially points of interest and 2D maps—dominate the literature because of their role in situational awareness, and that the gaps it identifies (limited multimodal feedback, few cross-discipline or modular designs, scarce comparative evaluation) are features of the literature itself.","pith_inferences":["The taxonomy's completeness is only as good as the search string; adding synonyms like 'wearable display' or 'see-through display' and searching forward and backward citations would likely surface additional papers and could reveal element types the 90-paper corpus missed.","The near-absence of haptic and auditory output (about 2 haptic and 10 auditory papers) may reflect publication bias toward visual AR rather than a true judgment that non-visual channels are ineffective, given that reviewed user feedback explicitly called for head-vibration alerts.","The paper's own evidence suggests the field is pre-paradigmatic: most papers propose systems, few evaluate them, so the reported 'gaps' are as much about evaluation culture as about the design space itself.","A practical extension would be to operationalize the taxonomy as a searchable database or generative design tool that first responder agencies could filter by facet values to request new capabilities."],"forward_implications":["Researchers can use the taxonomy to position new AR designs and identify unpopulated combinations, such as cross-discipline or multimodal configurations.","Developers get a checklist of interface elements with known design challenges, including visual clutter, cognitive load, and gesture reliability under stress.","The scarcity of comparative evaluations means effectiveness claims for most elements remain unproven, so future work should prioritize controlled comparisons against traditional methods.","Hands-free, context-aware, and modular interface design is identified as a priority direction for the field.","The pronounced growth in publications after 2016, tied to consumer XR hardware releases, suggests hardware cycles drive research attention in this domain."],"supporting_citations":[{"why":"Kitchenham and Charters provide the systematic review methodology the paper follows for search and selection.","marker":"[3]"},{"why":"Petticrew and Roberts define the PICOC framework used to structure the search terms.","marker":"[4]"},{"why":"Arksey and O'Malley supply the scoping review definition that justifies mapping breadth rather than assessing quality.","marker":"[5]"},{"why":"Peters et al. provide guidance for conducting systematic scoping reviews that the paper cites for its approach.","marker":"[7]"},{"why":"Endsley's situational awareness theory is the framework used to explain the function of the interface elements.","marker":"[99]"},{"why":"Ntoa et al. provide the DARLENE evaluation that supplies evidence object highlighting improves situational awareness under stress.","marker":"[57]"},{"why":"Zhang et al. supply first-responder user feedback on highlighting granularity and the necessity of head-vibration alerts.","marker":"[88]"},{"why":"Smets et al. support the claim that heading-up maps outperform north-up maps by reducing mental rotation.","marker":"[100]"},{"why":"Wilson and Wright flag usability issues with dot-based breadcrumbs, which the paper uses to suggest lines may be better.","marker":"[101]"}],"fun_headline_variants":["Six-facet taxonomy organizes 90 AR first-responder studies","90 studies, six facets: AR interface design space for first responders","AR for first responders: taxonomy of six facets from 90 studies","AR first-responder research misses multimodal feedback"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The map's completeness rests on the assumption that a keyword search of four English-language databases, restricted to terms in titles, abstracts, or keywords, retrieves a representative sample of the relevant literature; the paper acknowledges in Section 2.6 that papers not matching the search string were excluded.","fun_headline_variants_meta":{"raw":{"variants":["Six-facet taxonomy organizes 90 AR first-responder studies","90 studies, six facets: AR interface design space for first responders","AR for first responders: taxonomy of six facets from 90 studies","AR first-responder research misses multimodal feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000766,"raw_usage":{"total_tokens":3398,"prompt_tokens":945,"completion_tokens":2453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":2382}},"tokens_in":561,"tokens_out":2453,"duration_ms":16277,"temperature":1.0,"reasoning_tokens":2382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:53:18.801104+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that broadens the search—for example, adding Web of Science and PubMed, adding terms like 'wearable display' and 'see-through display,' and hand-searching the reference lists of the 90 included papers—would test the map's completeness. If such a search surfaced a sizeable set of relevant AR UI papers that do not fit into the six facets or 12 interface elements, or that substantially change the reported counts (e.g., many more haptic or auditory interfaces), the claim that Figure 4 maps the domain would be weakened.","supporting_citations":[{"cited_title":"Toward a theory of situation awareness in dynamic systems,","cited_arxiv_id":null,"evidence_quote":"Endsley's situational awareness theory is the framework used to explain the function of the interface elements."},{"cited_title":"A Mixed-Methods Approach for the Evaluation of Situational Awareness and User Experience with Augmented Reality Technologies,","cited_arxiv_id":null,"evidence_quote":"Ntoa et al. provide the DARLENE evaluation that supplies evidence object highlighting improves situational awareness under stress."},{"cited_title":"Ex- ploring the Design Space of Optical See-through AR Head- Mounted Displays to Support First Responders in the Field,","cited_arxiv_id":null,"evidence_quote":"Zhang et al. supply first-responder user feedback on highlighting granularity and the necessity of head-vibration alerts."},{"cited_title":"Effects of mobile map orientation and tactile feedback on navigation speed and situation awareness,","cited_arxiv_id":null,"evidence_quote":"Smets et al. support the claim that heading-up maps outperform north-up maps by reducing mental rotation."},{"cited_title":"Head-mounted display efficacy study to aid first responder indoor navigation,","cited_arxiv_id":null,"evidence_quote":"Wilson and Wright flag usability issues with dot-based breadcrumbs, which the paper uses to suggest lines may be better."}],"review_version":1}