{"id":"18a5c35c-c74f-4908-9689-7c9965760626","arxiv_id":"2507.14034","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Six interaction modes (HAM, HIC, HITP, HITL, HOTL, HOOTL) map human oversight levels to contingency factors for designing agentic AI in technical services.","lead":"The authors organize human-AI collaboration in technical services into six modes, from fully human-led to fully autonomous, and link each mode to factors like task complexity and operational risk. The framework is built from public documentation of AI features from Microsoft, Salesforce, and ServiceNow.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Empirical derivation is not supported by cited cases: HITP is an explicitly constructed 'composite use case' and HOOTL is asserted from a vendor blog with no documented absence of human oversight; the taxonomy should be framed as a design heuristic.","rationale":"I read the paper in good faith. It offers a coherent, practitioner-oriented framework, and the six modes are clearly described with process diagrams and vendor examples. However, the strongest claim explicitly invokes case-study research as the empirical basis for the taxonomy. That empirical derivation is the load-bearing condition: if modes are illustrated with hypothetical or asserted examples rather than documented deployments, the framework cannot claim to describe how human-AI collaboration is currently structured, only how it could be structured. The reader's weakest assumption concerned vendor-documentation accuracy; I partially agree, but I would sharpen the concern further. Within the paper itself, Section 4.3 confesses that HITP is a 'composite use case' based on platform potential, and Section 4.6 asserts HOOTL autonomy without quoting any source that rules out human oversight. These are internal evidence gaps, not merely worries about vendor exaggeration. A desk audit of the cited sources can settle whether the examples are real and whether the no-human-overview claims are textually supported. Because the paper can be salvaged by reframing the taxonomy as a provisional design heuristic and by removing unsupported empirical claims, the reader's conditional verdict remains appropriate rather than moving to reject.","tokens_in":12983,"tokens_out":5474,"duration_ms":60152,"concrete_test":"Audit the primary sources cited for the HITP and HOOTL modes (Winklix, 2025; ServiceNow, 2025a; O'Quinn, 2024; Rand Group, 2025; Sahye, 2025): extract every sentence concerning human approval, review, dispatch authorization, or oversight from these documents. If any source for the HITP workflow is not a named deployed system, or any source for HOOTL includes a human authorization or review step, then the two modes are not empirically instantiated and the central claim should be softened from 'empirically derived taxonomy' to 'conceptual design heuristic'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the six-mode taxonomy is an empirically derived, reusable framework rests on the cases in Sections 3 and 4 being real deployments accurately described by the cited sources. That condition fails internally. Section 4.3 labels the HITP example a \"composite use case\" and says it is \"exemplified by the potential of enterprise workflow platforms\" (ServiceNow), not an observed deployment; the only cited quotes are generic platform capabilities (Winklix, 2025; ServiceNow, 2025a), not a documented workflow with an actual human dispatch-approval step. Section 4.6 classifies Dynamics 365 Field Service as HOOTL on the basis of O'Quinn (2024) plus partner blogs, but the paper itself supplies no source sentence showing the end-to-end process actually runs without human review or approval; \"without human intervention/oversight\" is asserted in the authors' voice. If either mode is not instantiated by the evidence, the taxonomy's empirical basis is incomplete: one mode is hypothetical and the autonomy pole is inferred rather than documented. The framework may still be useful as a design heuristic, but the claim of case-study derivation overstates what the data show.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a six-mode taxonomy (HAM, HIC, HITP, HITL, HOTL, HOOTL) for human-AI collaboration in technical services, ranging from passive AI assistance to full autonomy. It claims to derive these modes from comparative case studies of three (or four) enterprise platforms — Microsoft Dynamics 365, Salesforce, and ServiceNow — using public vendor documentation collected between April and May 2025. The authors map each mode to contingency factors such as task complexity, operational risk, system reliability, and human operator state, and argue that the result is a reusable, technology-agnostic framework for selecting and designing human oversight in technical service systems.","tokens_in":13194,"tokens_out":5130,"duration_ms":54308,"significance":"If treated as a design heuristic rather than an empirically validated classification, the paper is a useful synthesis: it connects prior team design patterns (van Zoelen et al.) to concrete architectural primitives, offers a common vocabulary for practitioners, and makes falsifiable design conjectures, such as reserving HOOTL for low-risk, high-reliability tasks. The strengths are the clear conceptual organization, the process-level framing, and the explicit linkage between oversight modes and contingency factors. However, the evidence base is thin: the modes are illustrated with vendor marketing materials, one mode is explicitly a composite use case, and the autonomy pole is asserted rather than documented. The paper's value would increase substantially if these limitations were stated honestly and the framework labeled as a design heuristic rather than an empirical derivation.","major_comments":[{"comment":"The claim that the taxonomy is 'empirically derived' (Section 1) is not supported by the evidence. Section 4.3 explicitly introduces HITP as a 'composite use case' and says it is 'exemplified by the potential of enterprise workflow platforms,' not by an observed deployment; the cited sources (Winklix, 2025; ServiceNow, 2025a) describe platform capabilities rather than a running workflow with a human dispatch-approval step. Similarly, Section 4.6 asserts that the Dynamics 365 Field Service process 'operates without human oversight' in the authors' own voice; the cited O'Quinn (2024) and partner blogs do not document the absence of human review in a real deployment. Because two of the six modes rest on constructed or inferred examples, the empirical grounding of the taxonomy is incomplete. I recommend either reframing the contribution as a design heuristic grounded in illustrative vendor capabilities or adding evidence from actual deployments.","section":"4.3 and 4.6"},{"comment":"The evidence base consists entirely of official product pages, technical documentation, promotional videos, white papers, and corporate blogs. These are vendor-generated materials that may systematically overstate the autonomy of deployed systems, particularly regarding whether human oversight exists in practice. The paper does not discuss this bias or triangulate with independent sources, and Section 4.6 in particular relies on the absence of a documented human step as evidence for HOOTL. This is a load-bearing limitation for the central claim that the taxonomy describes real-world collaboration modes; it should be acknowledged and mitigated by framing the results as a synthesis of vendor capabilities rather than observed practice.","section":"3.2"},{"comment":"The case count is inconsistent: Table 1 lists three platform providers (Microsoft, Salesforce, ServiceNow), but Section 3.3 states that coding and comparison were performed 'across the four cases,' and Section 6 repeats 'four major technology providers.' This discrepancy is not merely typographical; it affects the claimed basis for the cross-case analysis. The authors should state the exact number of cases and align Table 1, Section 3.3, and Section 6.","section":"3.3 and 6"},{"comment":"The statement that 'True HOOTL automation requires exceptionally high system reliability (>95% accuracy)' is asserted without a citation, derivation, or operational definition. This numeric threshold is used to justify the HOOTL mode and is therefore load-bearing for the selection guidance. Either provide a source or empirical basis for the 95% figure, or remove the specific threshold and discuss reliability qualitatively.","section":"5.1"},{"comment":"Because the six modes are induced from the same three cases used to illustrate them, the taxonomy is at risk of circularity: the contingency factors in Section 5.1 are inferred from the same vendor materials that exemplify the modes. The paper should acknowledge this limitation explicitly and, ideally, test the framework on at least one independent case—for example, a non-vendor deployment or a different platform—before claiming it is 'reusable' and 'technology-agnostic.'","section":"3.3 and 5.1"}],"minor_comments":[{"comment":"The terminology for the first mode is inconsistent: the Abstract and Section 1 use 'Human-Augmented Model,' while Section 4.1 headings use 'Human-Augmentation-Mode.' Choose one form and use it throughout.","section":"1 and 4.1"},{"comment":"The reference to Hartikainen et al. contains corrupted HTML entities ('V&#228, &#228, n&#228, & Nen, K.') and should be corrected to the proper author names and title.","section":"2.3"},{"comment":"The sentence 'The AI can produce summaries of the customer query and prior and surface relevant knowledge articles' is grammatically incomplete; it appears to omit a noun such as 'conversations' after 'prior.'","section":"4.2"},{"comment":"The data collection period is stated as April–May 2025, but one cited source (D365 Community, 2025) is dated July 2025; please clarify whether the collection period extended beyond May or whether this reference was added during a later revision.","section":"3.2"},{"comment":"The figures are helpful, but the captions do not indicate whether they are the authors' own conceptual diagrams or adapted from vendor materials; add an explicit note in each caption or in the text.","section":"Figures 1–6"},{"comment":"The citation formatting in the sentence 'The HITL model (ServiceNow) (Roethof, 2025; ServiceNow, 2025b)' has duplicated parentheses; it should be cleaned up for readability.","section":"5.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a solid conceptual core and the taxonomy is coherent, but the methodological overreach in claiming empirical derivation is substantial. I believe the paper can be repaired through revision if the authors reframe the contribution as a design heuristic and address the specific evidence gaps; in its current form, the claim of case-study derivation would not withstand scrutiny."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is better than its packaging. The six-mode taxonomy (HAM, HIC, HITP, HITL, HOTL, HOOTL) is a reasonable, clearly explained extension of van Zoelen's three Team Design Patterns, and the contingency-factor mapping (task complexity, risk, reliability, operator state) gives practitioners something they can actually use to think about oversight levels. The writing is clean, the figures are helpful, and the authors have done their homework on the prior literature. If you need a shared vocabulary for agentic AI in technical services, this is a decent candidate.\n\nThe soft spots are real but concentrated in the empirical claims, not in the framework itself. Section 3.2 says the data are public vendor pages, blogs, and white papers—that is not case-study research on deployed systems, and the paper overreaches when it says the taxonomy is \"empirically derived.\" The HITP mode (Section 4.3) is explicitly labeled a \"composite use case\" built from platform capabilities, not an observed workflow. HOOTL (Section 4.6) is inferred from a Microsoft blog plus partner posts; no source sentence shows an end-to-end process running without any human review. The >95% reliability threshold appears in Section 5.1 with no citation. And the case count is inconsistent: three providers are selected in Section 3.2, but the conclusion says \"four major technology providers.\" These are fixable problems, but they need to be fixed.\n\nThe reader's stress-test note is on target about the empirical derivation. However, I do not think this sinks the paper. The taxonomy is internally coherent, and its value does not depend on the vendor cases being perfect instantiations. The authors should reframe the contribution as a design heuristic grounded in prior theory and illustrated by vendor materials, not as a case-study-derived empirical finding. They should also add a coding protocol or at least a transparent table of which sources support each mode, and they should drop or properly source the reliability threshold.\n\nFor whom? IS researchers working on human-AI teaming, and practitioners who need a structured way to choose an oversight model. It deserves a serious referee, but not acceptance as-is. I would send it to review with a clear request for major revision focusing on honest framing of the evidence base.\n\nRecommendation: engage with it, but push the authors to distinguish the analytical contribution from the illustrative cases.","headline":"A genuinely useful design heuristic for human-AI collaboration, but the empirical case-study framing overstates what the vendor documentation actually supports.","tokens_in":13725,"tokens_out":1354,"would_cite":true,"duration_ms":17491,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that human-AI collaboration in technical services can be structured along a six-mode autonomy spectrum, from passive AI assistance to full automation, with mode choice driven by task complexity, risk, reliability, and…","keywords":["human-AI collaboration","technical services","human-in-the-loop","autonomy spectrum","contingency factors","agentic AI","human-autonomy teaming","human-agent interaction"],"falsifier":"Observe a real Dynamics 365 Field Service predictive-maintenance deployment and check whether every work order is approved by a human before dispatch or whether supervisors monitor the queue. If any human approval or monitoring is required, the paper's HOOTL classification, and therefore the empirical anchor of the autonomy spectrum, is contradicted.","tokens_in":12736,"feed_emoji":"🤖","tokens_out":6802,"duration_ms":70806,"temperature":0.7,"pith_summary":"This paper tries to give technical service organizations a systematic way to decide how much control humans should keep over AI agents. It proposes a taxonomy of six interaction modes, spanning from passive AI assistance (Human-Augmented Mode) through mandatory human approval, built-in human steps, AI-initiated escalation, and discretionary supervision, to full automation (Human-Out-of-the-Loop). The authors argue that the right mode is not arbitrary but depends on four contingency factors: task complexity and novelty, safety and risk, system reliability and trust, and the human operator's workload and vigilance. If correct, managers and system designers gain a common vocabulary and a decision aid for trading off automation benefits against the risks of unreliable AI. The taxonomy is derived from comparing how Microsoft, Salesforce, and ServiceNow describe their technical service AI products.","feed_headline":"Six-mode taxonomy links AI oversight to task risk and complexity","feed_subtitle":"Technical service teams can use it to pick the right autonomy level before deploying agentic AI.","key_machinery":"The carrying object is the six-mode taxonomy itself: HAM, HIC, HITP, HITL, HOTL, and HOOTL, defined along a spectrum of AI autonomy. Each mode is distinguished by two questions: who owns each activity, and what triggers human involvement, whether mandatory approval, a fixed workflow step, AI-initiated escalation, discretionary supervision, or nothing. The taxonomy is anchored to a standard service process framework and to four contingency factors, namely task complexity and novelty, safety and criticality, system reliability and trust, and human operator state, so that a mode is not just a label but a design configuration. The case analysis of vendor documentation provides the empirical instantiations, such as Salesforce Agentforce for HIC and ServiceNow virtual agents for HITL.","core_discovery":"The central claim is that every human-AI collaboration in technical services can be described by one of six interaction modes that form an autonomy spectrum. In HAM the AI only augments a human who does all work; in HIC the AI drafts solutions but a human must approve before anything is sent; in HITP the AI runs a workflow that stops at pre-engineered human steps; in HITL the AI operates until its confidence drops and then escalates to a human; in HOTL the AI runs end-to-end while a human supervisor may intervene at will; and in HOOTL the AI completes the whole process with no human involvement. The paper maps these modes onto the technical service process of receipt, diagnosis, solution, approval, communication, and closure, and links them to contingency factors that should drive selection. The proposed payoff is a reusable, technology-agnostic framework that turns the choice of human oversight level from an ad hoc decision into a structured design step.","pith_inferences":["A testable extension is to treat the contingency factors as a decision rule and validate whether organizations that follow it achieve fewer failures than those that choose modes ad hoc.","The same six-mode structure could be applied outside technical services, including customer support, healthcare triage, and logistics, since the triggers that define the modes are domain-neutral; this extrapolation goes beyond what the paper argues.","If vendor materials overstate automation, real deployments may sit at a lower-autonomy mode than the classification suggests, and a deployment-level audit would reveal mismatches that could refine the taxonomy.","A natural next design step, which the paper names only as future research, is a dynamic mode that shifts between HITL and HOTL in real time based on confidence scores and operator workload."],"forward_implications":["Technical service teams can map an existing or planned process onto the six modes and use the contingency factors to justify a specific oversight level before deployment.","The same six labels apply across vendors and platforms, giving design conversations a common vocabulary that does not depend on a particular product's terminology.","High-risk, safety-critical services will tend to stay in HIC or HAM, while low-risk, repetitive, well-understood tasks can move to HOTL or HOOTL when measured reliability is high.","The taxonomy gives evaluators a fixed set of configurations for comparing error rates, completion time, and operator satisfaction in empirical studies.","Organizations can use the framework to run structured risk assessments, connecting a chosen mode to task complexity and operational risk rather than relying on vendor promises."],"supporting_citations":[{"why":"It defines the three Team Design Patterns that the taxonomy generalizes into six modes.","marker":"(Van Zoelen et al., 2023)"},{"why":"It supplies the two-axis distinction between automation and technical autonomy used to separate the six modes.","marker":"(Simmler & Frischknecht, 2021)"},{"why":"It provides the large-scale productivity evidence that motivates augmenting human agents with AI rather than fully automating.","marker":"(Brynjolfsson et al., 2023)"},{"why":"It supplies the standard technical service process stages onto which the modes are mapped.","marker":"(Wulf & Winkler, 2020)"},{"why":"It is the market survey used to choose leading platforms for the comparative case study.","marker":"(Gartner, Inc., 2025)"},{"why":"It is the vendor documentation of Dynamics 365 Customer Service used to instantiate the Human-Augmented Mode.","marker":"(Microsoft, 2025)"},{"why":"It is the vendor documentation of Agentforce service replies used to instantiate the Human-in-Command mandatory review gate.","marker":"(Salesforce, 2025b)"},{"why":"It is the vendor documentation of NLU confidence thresholds used to instantiate the Human-in-the-Loop escalation mechanism.","marker":"(ServiceNow, 2025b)"},{"why":"It is the Microsoft field service blog used to instantiate the Human-Out-of-the-Loop predictive maintenance scenario.","marker":"(O’Quinn, 2024)"}],"fun_headline_variants":["Six modes map human-AI control in technical services","Pick the right AI oversight level with six-mode framework","Human-AI teamwork taxonomy for technical service autonomy","Six interaction modes span full AI to human-only control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that public vendor documentation from April to May 2025 accurately describes how these AI systems actually behave in deployed technical service settings, specifically that Dynamics 365 Field Service operates without human oversight.","fun_headline_variants_meta":{"raw":{"variants":["Six modes map human-AI control in technical services","Pick the right AI oversight level with six-mode framework","Human-AI teamwork taxonomy for technical service autonomy","Six interaction modes span full AI to human-only control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000653,"raw_usage":{"total_tokens":3046,"prompt_tokens":1053,"completion_tokens":1993,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":1931}},"tokens_in":669,"tokens_out":1993,"duration_ms":17926,"temperature":1.0,"reasoning_tokens":1931,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:10:37.042358+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Observe a real Dynamics 365 Field Service predictive-maintenance deployment and check whether every work order is approved by a human before dispatch or whether supervisors monitor the queue. If any human approval or monitoring is required, the paper's HOOTL classification, and therefore the empirical anchor of the autonomy spectrum, is contradicted.","supporting_citations":[],"review_version":1}