{"id":"09035247-2c5a-4a0f-aa08-67287b29cc49","arxiv_id":"2412.14232","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that many 'human-in-the-loop' systems are actually 'AI-in-the-loop' because the human controls the decision while AI assists, and that this difference should guide evaluation.","lead":"This paper renames many 'human-in-the-loop' AI systems as 'AI-in-the-loop' systems, arguing that the human, not the AI, is the true decision-maker in these cases. It asks researchers to separate the two categories and to judge AI systems by how well they support human goals, not just by model accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The HIL/AI2L taxonomy rests on an unoperationalized notion of control; the paper's own Discussion concedes the domains are 'nested,' making the binary classification depend on analysis granularity and undermining the prescriptive evaluation claims.","rationale":"The reader's weakest_assumption correctly identifies the lack of an operational definition of control as the load-bearing weakness, and the paper's own 'nested' concession in the Discussion supports this. My analysis agrees: the central claim that AI2L systems are systematically mislabeled as HIL, and that evaluation should be human-centric for AI2L, depends on being able to classify systems reliably. The concrete test would settle whether the taxonomy is stable across reasonable formalizations of control. I find no grounds to reject the paper outright; it is a position paper with a plausible thesis, and the conditional verdict is appropriate. The paper has independent value in drawing attention to evaluation mismatches and in referencing prior work (e.g., van Amsterdam et al., Selbst et al.), but it does not provide a crisp criterion for applying its labels. The honest outcome is to keep the CONDITIONAL verdict, requiring an operationalization before the prescriptive claims can be adopted. I see no need to escalate to REJECT, as the contribution is conceptual rather than empirical, and the limitations are acknowledged in the text rather than hidden. The lack of formal verification and the absence of systematic literature review are secondary to the definitional problem, which undermines the applicability of the taxonomy to concrete systems.","tokens_in":10003,"tokens_out":2433,"duration_ms":24279,"concrete_test":"Take Table 1's 'Fraud detection' row and formalize decision rights at three levels: (1) which events trigger the AI's flagging algorithm, (2) whether human confirmation is a mandatory gate before any action is taken, and (3) who sets the threshold that defines 'suspicious.' Classify the same system using the paper's Figure 1 definition ('AI systems drive inference and decision-making' vs. 'humans make the ultimate decisions') separately at each level. If the classification flips across levels (e.g., HIL at level 2 but AI2L at level 3), the taxonomy is unstable. As a complementary empirical check, present a set of 20 real-world HCI/AI systems covering different delegation patterns to a panel of researchers familiar with the paper, ask them to label each as HIL or AI2L using only the paper's definition, and compute inter-annotator agreement (e.g., Cohen's kappa).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that systems can be separated by who holds 'decision-making authority and control' (Introduction, Figure 1 caption), and that this separation should drive distinct evaluation regimes. However, 'control' is never defined operationally. In the Discussion, the paper states: 'the appropriate problem domains for the HIL and AI2L systems are typically not separated but nested... zooming in on a domain would result in a HIL problem and zooming out would give us an AI 2L problem.' This concession is decisive: if the same system is HIL when described at one granularity and AI2L at another, then the classification is not an intrinsic property of the system, and the paper's categorical prescriptions (e.g., 'evaluations of these systems are human-centric' for AI2L versus AI-centered metrics for HIL) lack a stable foundation. The two introductory examples depend on an intuitive reading: the recommender system is called HIL because 'the AI agent optimizes its internal function... and decides,' while the physician example is AI2L because 'the human is in control of the full system.' But no decision procedure is given for cases where the human's role is a probabilistic gate, where the human sets incentives that shape AI decisions, or where control is distributed across roles (e.g., Table 1's fraud detection, where AI flags and 'humans confirm'). Without such a procedure, the classification cannot be reliably applied, and the paper's claim that 'existing evaluation methods often overemphasize the machine component's performance' cannot be tested against specific systems. This is not a disagreement with the blue-sky vision; it is a missing precondition for the paper's core contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This blue-sky position paper argues that many systems commonly called human-in-the-loop (HIL) are in fact AI-in-the-loop (AI²L) systems: the human retains decision-making authority and the AI serves as a supporting component, whereas in true HIL systems the AI is the decision-maker and the human supplies labels, feedback, or advice. The authors contend that existing evaluation practice overemphasizes machine-learning metrics such as accuracy and precision, which is appropriate for HIL systems but inadequate for AI²L systems, and they advocate human-centric evaluation that accounts for the human expert's role, the human-AI interaction, and end goals such as health outcomes. The paper motivates the distinction with two introductory examples, a comparison figure, a table of domain examples, and a discussion of differences in control, bias, and evaluation.","tokens_in":10212,"tokens_out":4358,"duration_ms":41115,"significance":"If the proposed distinction were made precise and well-grounded, it would be a genuinely useful corrective: the field does often conflate systems in which the AI is the primary actor with systems in which it is a subordinate assistant, and the paper's connection to abstraction errors and human-centric evaluation is timely. The strengths of the paper as a position piece are its clear Figure 1, the concrete Table 1 examples across six domains, and its alignment with prior work on explainable/advisable AI and on the limitations of static benchmarks. However, the paper's central contribution is a taxonomy, and the taxonomy currently rests on an unoperationalized notion of control; the manuscript itself concedes that the categories are nested and granularity-dependent. This limits the paper's actionable impact and, as presented, weakens the prescriptive evaluation claims. The paper is a reasonable blue-sky argument but needs a sharper formalization before it can serve as a reliable framework for practitioners.","major_comments":[{"comment":"The paper's own Discussion states that 'the appropriate problem domains for the HIL and AI2L systems are typically not separated but nested' and that 'zooming in on a domain would result in a HIL problem and zooming out would give us an AI2L problem.' This makes the HIL/AI2L classification a function of analysis granularity rather than an intrinsic property of the system. Since the paper's central prescriptions—AI-centered metrics for HIL versus human-centric evaluation for AI²L—depend on a stable classification, this concession undermines the load-bearing claim. Please provide a decision procedure for selecting the level of abstraction at which the classification is made, or state an explicit rule for when a system should be evaluated as HIL versus AI²L despite the nested relationship.","section":"Discussion (nested domains paragraph)"},{"comment":"The central notion of 'decision-making authority and control' is never defined operationally. The two introductory examples are classified by an intuitive reading of who 'decides,' and Table 1's fraud-detection row is classified as 'Automate (HIL)' even though the description says 'AI analyzes the data and flags suspicious activities; humans confirm'—a configuration that could equally be read as human-in-control if confirmation is decisive. Without explicit criteria (e.g., who has veto power, who bears final accountability, whether the system can operate without the human), the taxonomy cannot be reliably applied by other researchers, and the claim that a system should be evaluated human-centrically or AI-centrically is not testable. Please add operational criteria for control and apply them consistently to all Table 1 examples.","section":"Introduction and Figure 1 caption"},{"comment":"The paper asserts that 'existing evaluation methods often overemphasize the machine (learning) component's performance, neglecting the human expert's critical role,' but provides no survey, taxonomy of metrics, or quantitative evidence to support this empirical premise. This premise motivates the entire AI²L evaluation agenda, so it is load-bearing rather than a decorative observation. Please support it with a structured review of typical evaluation protocols in the cited application areas, or at minimum with several concrete cases where a published 'HIL' system is evaluated using AI-centered metrics that obscure the human contribution. Without such support, the paper's central diagnostic claim remains an assertion.","section":"Abstract and the AI-in-the-loop evaluation paragraph"}],"minor_comments":[{"comment":"Using 'HIL' to denote a system in which the AI is in control and 'AI²L' to denote a system in which the human is in control is counterintuitive and will likely confuse readers, since 'human-in-the-loop' commonly implies human oversight. Please add an explicit note contrasting the paper's usage with the common reading, or consider relabeling the categories (e.g., 'AI-centered' versus 'human-centered').","section":"Throughout (terminology)"},{"comment":"The sentence 'calling them HIL would be a misnomer, as they are quite the opposite, namely AI-in-the-loop systems, where the human is in control of the system' is self-contradictory on its face: if the human is in control, then a human is literally in the loop. Clarify that 'AI-in-the-loop' is meant to describe the AI's subordinate role rather than who holds control.","section":"Abstract"},{"comment":"The heading 'Collaborate (AI2L)?' contains a stray question mark, and the notation is inconsistent between 'AI²L' and 'AI 2L' across the paper; please standardize the symbol and its spacing.","section":"Table 1"},{"comment":"The reference 'Wang, G. 2019' is incomplete, with no title or venue; please supply the full citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a blue-sky position piece, so the lack of an empirical survey is not itself disqualifying, but the control-based taxonomy must be operationalized and the nestedness concession addressed for the paper to be more than a terminological proposal. The authors may also wish to engage with the human-factors literature on levels of automation and human-machine delegation, which directly addresses the granularity issue the paper acknowledges. The paper is within scope for a workshop/blue-sky venue but would need substantial revision to be a definitive reference."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a clearly written blue-sky position paper with a useful name for a real distinction. The central idea — separate \"AI in control, human advises\" (HIL) from \"human in control, AI advises\" (AI2L) and evaluate them differently — lands well, and the examples in Table 1 help make it concrete. The authors are also honest: they cite Kambhampati, Sreedharan, and van Amsterdam, so they don't pretend the underlying thoughts are brand new. The contribution is the packaging and the push to take evaluation seriously.\n\nThe soft spots are real but proportionate. The stress test is right that \"control\" is never defined operationally. The recommender vs. physician example works only on an intuitive reading, and the Discussion concedes that the domains are \"nested\" — zoom in and you get HIL, zoom out and you get AI2L. That makes the binary classification granularity-dependent, which undercuts the categorical evaluation prescriptions. A second soft spot: the claim that existing evaluations \"overemphasize the machine component\" is asserted, not demonstrated. There is no survey, no metric taxonomy, no quantitative evidence. For a blue-sky paper that's not disqualifying, but it limits persuasiveness.\n\nI'd send this to peer review. It's the kind of paper that generates useful discussion at a workshop or in a blue-sky track. The authors should be asked to either operationalize control (e.g., who has veto or final authority over consequential actions) or soften the taxonomy into a spectrum with prototypes. As is, I'd read it as a call-to-arms rather than a settled framework. Who benefits: HCI researchers, evaluation-methodology people, and anyone designing decision-support systems. I don't think I'd cite it in my own work in the next year, but I'd point students to it as a clear statement of the distinction.","headline":"A clear, honest blue-sky paper with a useful \"AI2L\" label; the binary taxonomy is underspecified on control, but it deserves a serious referee rather than a desk reject.","tokens_in":10886,"tokens_out":1725,"would_cite":false,"duration_ms":16411,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that many human-in-the-loop systems are misnamed: when the human holds final decision authority and the AI assists, the system is AI-in-the-loop, and evaluation must be human-centered.","keywords":["human-in-the-loop","AI-in-the-loop","human-AI collaboration","evaluation metrics","decision authority","interactive machine learning","human-centered evaluation","automation vs collaboration"],"falsifier":"Run a reliability study in which two independent coders apply the paper's definitions to the systems listed in its examples; any substantial disagreement on the HIL versus $AI^{2}$L label would show that \"who is in control\" is not an operational test.","tokens_in":9751,"feed_emoji":"🔄","tokens_out":6851,"duration_ms":55848,"temperature":0.7,"pith_summary":"Many systems described as human-in-the-loop (HIL) put the machine in charge and use people as labelers or advisors; this paper argues that a large share of them should instead be called AI-in-the-loop ($AI^{2}$L), where the human expert holds decision authority and the AI assists. The paper claims this mislabeling is not cosmetic: it leads to evaluating the wrong thing, measuring model accuracy when the relevant outcome is the quality of the human's decision, and it can produce abstraction errors during design and deployment. The authors propose that designers first decide whether they are automating a well-defined subtask (HIL) or collaborating inside a human decision process ($AI^{2}$L), and then adopt evaluation regimes that fit the actual control structure. If right, the paper gives the field a sharper vocabulary and a reason to make human-centered metrics, including ablation studies, interpretability, fairness, and end-goal impact, standard for collaborative systems.","feed_headline":"Call them AI-in-the-loop, not human-in-the-loop","feed_subtitle":"When the human makes the final call and AI assists, accuracy alone is the wrong test; measure outcomes instead.","key_machinery":"The load-bearing object is a two-way taxonomy built on the locus of decision-making authority and control. HIL (automate) is defined as an autonomous AI agent that may seek human input; $AI^{2}$L (collaborate) is defined as an intervention in a human decision-making process, where the AI presents synthesized information, possible actions, and consequences and the human chooses. The distinction does the argument's work: it predicts different sources of bias (model and data bias versus human interpretation bias), different trust problems (human-teacher credibility versus system transparency), and different evaluation requirements (AI-centered metrics versus human-centered, end-goal-aligned metrics). The paper's key move is to make this control-based classification the first step in system design and to show, through its table of examples, that many real subtasks in medicine, driving, logistics, manufacturing, finance, and education sort cleanly into either automate (HIL) or collaborate ($AI^{2}$L).","core_discovery":"The central claim is that 'human-in-the-loop' and 'AI-in-the-loop' name opposite control configurations. In a true HIL system the AI is the decision-maker and the human supplies corrections, labels, advice, or adversarial input to steer it; in an $AI^{2}$L system the human is the decision-maker and the AI is an assistive component that summarizes evidence, proposes options, and flags risks. Because the two configurations differ in who holds authority, they also differ in where bias enters and in what counts as success. The paper therefore contends that treating $AI^{2}$L systems as HIL produces wrong evaluation (accuracy-centered rather than human-outcome-centered), wrong trust modeling (credibility of the human teacher rather than transparency of the system), and abstraction errors in deployment. Its constructive proposal is to classify systems by control before design, and for $AI^{2}$L systems to evaluate with ablation, interpretability, fairness, and impact on the human's actual decision.","pith_inferences":["If the control-based taxonomy is right, a natural test is to replace the binary with a graded \"control profile\" measuring how much discretion the human actually exercises per subtask; the paper's own admission that domains are nested suggests the binary is a heuristic, not a law.","A concrete extension: for AI^2L systems, model evaluation should be supplemented by decision-level causal analysis that estimates the AI's marginal effect on human choices, not just its predictive accuracy.","The paper's framing implies that benchmark suites for human-AI collaboration should record the human's final decision and downstream outcome, not only the AI suggestion; existing datasets that log only model outputs would be insufficient for AI^2L evaluation.","Applied to foundation models, the paper's argument points to a research agenda in which models are trained to represent user goals and uncertainty rather than merely to follow instructions, because only then do they shift from HIL-style reactiveness to AI^2L-style collaboration."],"forward_implications":["Adopting the distinction means moving AI^2L evaluation away from accuracy, precision, and recall as the primary yardstick and toward human-centered metrics such as ablation studies, interpretability, fairness, and measured impact on the human's decision outcome.","Designers who currently build \"HIL\" systems for collaborative tasks would be forced to ask up front whether they are automating a well-defined subproblem (HIL) or intervening in an open-ended human decision process (AI^2L), changing what gets abstracted away.","The paper claims the framework transfers beyond supervised learning to reinforcement learning, planning, continual learning, and foundation models; in particular, large language models with thumbs-up or thumbs-down feedback are HIL-style oversight systems, not true collaborators, unless they acquire a theory of mind.","A correct HIL or AI^2L classification would reduce abstraction errors of the kind that occur when a complex, context-dependent domain is treated as a neat automation problem."],"supporting_citations":[{"why":"Defines active learning and the labeler-oracle view of humans, the HIL paradigm the paper contrasts with AI^2L.","marker":"Settles 2009"},{"why":"Introduces interactive machine learning, the shared-control lineage from which AI^2L is drawn.","marker":"Fails and Olsen Jr 2003"},{"why":"Names the broader role of humans in interactive machine learning and motivates human-centered evaluation that AI^2L extends.","marker":"Amershi et al. 2014"},{"why":"Supports the claim that evaluating by performance alone suits HIL but not AI^2L, and that clinical AI should be judged by causal impact on patient outcomes.","marker":"van Amsterdam et al. 2024"},{"why":"Supplies the abstraction-error concept the paper uses to argue that mislabeled HIL systems fail sociotechnical evaluation and deployment.","marker":"Selbst et al. 2019"},{"why":"Grounds the related vision of human-AI symbiosis that AI^2L extends toward user- and population-specific evaluation.","marker":"Kambhampati et al. 2022"},{"why":"Surveys human-in-the-loop machine learning as it is commonly understood, providing the target of the paper's relabeling argument.","marker":"Mosqueira-Rey et al. 2023"}],"fun_headline_variants":["When AI suggests and humans decide, it's AI-in-the-loop","Judge AI-in-the-loop by human outcomes, not model accuracy","Call it AI-in-the-loop when the human holds the reins","AI-in-the-loop: when the human is the boss","Not all loops are human-in-the-loop"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy assumes that a clear line can be drawn between \"the human is in control\" and \"the AI is in control\" in real systems, even though control is often shared, shifting, or context-dependent.","fun_headline_variants_meta":{"raw":{"variants":["When AI suggests and humans decide, it's AI-in-the-loop","Judge AI-in-the-loop by human outcomes, not model accuracy","Call it AI-in-the-loop when the human holds the reins","AI-in-the-loop: when the human is the boss","Not all loops are human-in-the-loop"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00084,"raw_usage":{"total_tokens":3640,"prompt_tokens":904,"completion_tokens":2736,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":2654}},"tokens_in":520,"tokens_out":2736,"duration_ms":19099,"temperature":1.0,"reasoning_tokens":2654,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:26:17.222965+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a reliability study in which two independent coders apply the paper's definitions to the systems listed in its examples; any substantial disagreement on the HIL versus $AI^{2}$L label would show that \"who is in control\" is not an operational test.","supporting_citations":[{"cited_title":"A.; and Olsen Jr, D","cited_arxiv_id":null,"evidence_quote":"Introduces interactive machine learning, the shared-control lineage from which AI^2L is drawn."},{"cited_title":"B.; and Kulesza, T","cited_arxiv_id":null,"evidence_quote":"Names the broader role of humans in interactive machine learning and motivates human-centered evaluation that AI^2L extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that evaluating by performance alone suits HIL but not AI^2L, and that clinical AI should be judged by causal impact on patient outcomes."},{"cited_title":"D.; danah boyd; Friedler, S","cited_arxiv_id":null,"evidence_quote":"Supplies the abstraction-error concept the paper uses to argue that mislabeled HIL systems fail sociotechnical evaluation and deployment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the related vision of human-AI symbiosis that AI^2L extends toward user- and population-specific evaluation."}],"review_version":1}