{"id":"c891086e-b0ad-48d1-99d5-e73c5a8a8826","arxiv_id":"2607.09734","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Proactive robots apply individual-targeting logic in 60.3% of group HRI studies, leaving detection and negotiation of entry into pre-formed groups unaddressed.","lead":"A systematic review of 63 proactive HRI studies finds that robots treat people in groups as independent targets 60% of the time, and almost never model how to join a pre-formed group. This reframes group HRI as a distinct design problem and flags group entry as an open challenge for public robots.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.5","headline":"The entry-gap claim rests on a dual-coding and multi-context counting scheme that can reclassify the same paper as both ITA and GAA, so the 60.3% and \"unaddressed\" figures are sensitive to how dual-coded and multi-context papers are tallied.","rationale":"The Reader correctly flags unquantified theme-level reliability and the inclusive GAA threshold as the main fragility. That is real, but the more load-bearing point for the paper's distinctive claim is not generic coding noise; it is how dual-coding and multi-context counting specifically underwrite the entry-gap assertion that \"remains unaddressed across the corpus.\" The screening kappa (0.812) is solid; the OSF codebook helps; the qualitative reframing of group entry as a distinct design problem is well motivated by Goffman/Kendon and by the failure-mode chain. Independent re-coding of the dual/multi-context subset is the single check that would settle whether the quantitative absolute language is stable. Until that is done, CONDITIONAL remains the right verdict: useful structural diagnosis, but the 60.3% and \"unaddressed\" headlines should not be treated as definitive without the re-tally. This is a partial agreement with the Reader—same family of concern (coding stability), different focal point (entry-gap counting vs. overall theme reliability).","tokens_in":14752,"tokens_out":789,"duration_ms":8542,"concrete_test":"Independently re-code the 14 multi-context papers and the four dual-coded papers ([42],[43],[44],[15]) for primary interaction context and for whether any GAA technique is applied before contact initiation (not only after embedding). Recompute Table II Group Entry row and the ITA total excluding dual-codes from the ITA side; if Group Entry GAA rises above 1–2 papers that truly address pre-contact group openness, or if ITA falls below ~50%, the absolute \"unaddressed across the corpus\" claim and the 60.3% figure both weaken.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's strongest claim is not merely that ITA is common, but that group-aware behaviour is almost entirely post-entry and that \"how a robot should detect and negotiate entry into a pre-formed group before initiating contact remains unaddressed across the corpus\" (Abstract; §V.B; Table II note). That claim depends on two coding decisions that are only partially transparent: (1) four papers are dual-coded as both ITA and GAA (Table I note: [42],[43],[44],[15]), and (2) 14 papers span multiple interaction contexts, so row totals exceed N=63 (Table II note). For Group Entry specifically, the authors report 24 papers (22 ITA, 4 GAA) and then qualify that the four GAA cases apply group-aware techniques only after interaction has begun (or only to spatial conduct via F-formation in Joosse et al. [45]), not to pre-contact openness detection. Because dual-coding and multi-context assignment are done by consensus without a theme-level reliability statistic (§III.D; Limitations), a modest reclassification of a few dual-coded or multi-context papers could shrink the apparent entry gap or change whether the \"unaddressed\" claim remains absolute. The 60.3% ITA prevalence is likewise an aggregate of Unreflective+Critical+Transitional that includes dual-coded papers, so the headline percentage is not a pure partition. The qualitative diagnosis (entry is harder and understudied) can survive, but the quantitative framing of the gap as corpus-wide and absolute is the load-bearing, least secure step.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"This systematic review of 63 proactive HRI studies (2000–2025) argues that robots in multi-human settings often treat co-present people as independent engagement targets—the Individual-Targeting Assumption (ITA)—rather than as a relational social unit. The authors report ITA in 60.3% of the corpus (Unreflective, Critical, and Transitional forms), catalogue Group-Aware Approaches (GAA) that model relational properties, and map three interaction contexts (group entry, facilitation, spatial conduct). They document three recurring failure modes (engagement misdetection, social ratification blindness, bystander neglect) and conclude that group-aware techniques appear almost entirely after the robot is already embedded, leaving pre-contact detection and negotiation of entry into a pre-formed group unaddressed. Proactive group HRI is thereby reframed as a qualitatively distinct design problem whose critical open challenge is the entry phase.","tokens_in":15229,"tokens_out":1567,"duration_ms":18302,"significance":"If the prevalence and entry-gap claims hold, the paper supplies a useful conceptual vocabulary (ITA vs GAA) and a concrete research agenda for a setting that is common in public deployments but still under-theorized relative to dyadic proactive HRI. Strengths include a transparent multi-database search, explicit inclusion/exclusion criteria, high inter-rater agreement on study inclusion (κ=0.812), and clear summary tables that make the coding scheme inspectable. The synthesis of failure modes and the catalogue of group-aware techniques (participation management, structure sensing, social-affective regulation, trust modelling) will be of practical value to designers. The contribution is primarily diagnostic and agenda-setting rather than a new algorithm or theory; its impact depends on the stability of the coding-derived percentages and on whether the absolute “unaddressed across the corpus” claim for pre-entry group negotiation survives closer scrutiny of dual-coded and multi-context papers.","major_comments":[{"comment":"Table I note and Table II (Group Entry row and note): Four papers are dual-coded as both ITA and GAA, and 14 papers span multiple interaction contexts so row totals exceed N=63. The headline ITA rate (60.3%) aggregates Unreflective+Critical+Transitional and includes dual-coded items; the absolute claim that pre-contact entry negotiation is “unaddressed across the corpus” (Abstract; §V.B) rests on reinterpreting the four GAA-in-entry cases as applying group-aware methods only post-initiation or only to spatial conduct (Joosse et al.). Because dual-coding and multi-context assignment are consensus-based without a reported theme-level reliability statistic, modest reclassification of a few papers could shrink the apparent gap or make the “unaddressed” claim non-absolute. Please (i) report a pure partition (e.g., exclusive ITA / exclusive GAA / dual) and sensitivity of the 60.3% and entry co","section":"Table I; Table II; §V.B"},{"comment":"§III.D and Limitations: Inclusion screening reports Cohen’s κ=0.812, but theme codes (ITA spectrum, GAA, three contexts, three failure modes, techniques) were finalized by iterative consensus with no formal inter-rater reliability on the theme-level codes. For a systematic review whose central quantitative claims are prevalence and failure counts, this is a load-bearing methodological gap. At minimum, report double-coding of a substantial subsample of the final 63 papers with κ (or equivalent) for the main theme codes, or provide a transparent audit trail (codebook excerpts and decision rules for dual-coding and multi-context assignment) so readers can assess stability of the 60.3%, 22/24 entry-ITA, and failure-mode n’s.","section":"§III.D; Limitations"},{"comment":"Table III and §IV (Failure Mode theme): Failure modes are coded when breakdowns “could be interpreted as related to individual targeting,” and the authors note that this required interpretive judgment; several GAA papers are dual-coded because they document failures as baselines. The causal-chain narrative in §V.C (misdetection → ratification blindness → bystander neglect) is plausible but not independently measured in the corpus. Please separate (a) failures reported in ITA-only papers from (b) failures used as motivating baselines in GAA papers, and avoid implying a measured causal sequence unless the source papers support it. Clarify whether n=6/9/18 are unique papers or applications, and whether papers can contribute to multiple modes without double-counting in the “most common” claim.","section":"Table III; §IV; §V.C"}],"minor_comments":[{"comment":"Abstract and Introduction state the review covers 2000–2025, while §I also says papers were reviewed from 2009–2025 because none met criteria before 2009. Align the date range wording so readers are not led to expect pre-2009 included studies.","section":"Abstract; §I"},{"comment":"Table II Group Entry: text says “22 of 25 papers” in §V.B but Table II lists n=24 (38.1%) for Group Entry. Reconcile 24 vs 25.","section":"§V.B; Table II"},{"comment":"Table IV note: “group turn-taking management (n = 19)” in the Discussion text vs n=18 in Table IV. Align counts.","section":"Table IV; §V.D"},{"comment":"Fig. 1 and Fig. 3 are described but would benefit from explicit callouts in the Findings when ITA/GAA and the three failure modes are first formalized, so the figures are not only introductory.","section":"§IV"},{"comment":"Minor typography: “V ázquez” / “V´azquez” spacing; “andbystander” missing space in Abstract; “engagement misdetection,social ratification” missing spaces in Abstract. Clean for production.","section":"Abstract; References"},{"comment":"OSF link is given for corpus and codebook; ensure the deposit is anonymized for review if required by the venue, and state whether the full dual-coding matrix is included.","section":"§III.D"}],"recommendation":"major_revision","confidential_remarks":"Solid agenda-setting review for cs.HC / HRI; novelty is in naming ITA and quantifying the entry gap rather than in new methods. The absolute “unaddressed” claim is the main risk for overclaim; if the authors supply sensitivity analyses and theme-level reliability, the paper is close to a useful contribution. Scope fit is good for a systematic-review track. No concerns about misconduct or citation gaming beyond normal self-citation of related team work in the corpus."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean, useful systematic review. The new piece is not \"groups matter\" (Sebo et al. and van Den Broek already said that). It is the named Individual-Targeting Assumption, the 60.3% prevalence over a purpose-built 63-paper corpus, the three failure modes, and the sharp claim that group-aware work almost never covers pre-contact entry into a pre-formed group.\n\nWhat they did well: transparent PRISMA-style screening, high inclusion kappa (0.812), explicit criteria, OSF codebook, and tables that separate Unreflective/Critical/Transitional ITA from GAA and map techniques to facilitation vs entry vs spatial conduct. The causal chain they sketch (misdetection → ratification blindness → bystander neglect) is readable and matches the cited field studies. The entry-gap argument is the real contribution: facilitation has tools; entry does not, and the information problem is harder because the robot has only a short observation window.\n\nSoft spots, in proportion. Theme-level coding is consensus-only with no reliability statistic; failure-mode assignment is interpretive by their own admission. Four papers are dual-coded ITA+GAA and 14 span contexts, so the 60.3% and the absolute \"unaddressed across the corpus\" line are sensitive to how those papers are tallied. The stress-test note is right that a modest reclassification could shrink the quantitative gap; the qualitative diagnosis (entry is understudied and harder) still holds. No temporal trend analysis, mostly lab studies, inclusive GAA threshold. None of that sinks the paper.\n\nWho it is for: anyone designing proactive public robots or multiparty HRI evaluation. Cite it for the construct and the open problem, not as a definitive prevalence law. I would send it to peer review; it is important enough and grounded enough to deserve referee time, with a request to tighten dual-coding transparency and theme reliability. Worth reading and citing.","headline":"Solid systematic review that names and quantifies ITA and isolates a real group-entry gap; the 60.3% and \"unaddressed\" claims are a bit soft on dual-coding and theme reliability, but the diagnosis is useful and worth engaging.","tokens_in":15790,"tokens_out":512,"would_cite":true,"duration_ms":5593,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Proactive robots treat people in groups as independent targets rather than social units, and how they should join a pre-formed group remains unaddressed across 63 studies.","keywords":["proactive HRI","human groups","Individual-Targeting Assumption","group entry","multiparty interaction","social robots","engagement","bystander neglect"],"falsifier":"Independent re-coding of the same 63 papers with the published codebook, or an expanded search that finds multiple proactive systems that explicitly model group openness and negotiate entry before first contact; a large share of group-aware entry designs would overturn the claim that entry is unaddressed.","tokens_in":15623,"feed_emoji":"🤖","tokens_out":1002,"duration_ms":19196,"temperature":0.7,"pith_summary":"Proactive robots in public places almost always meet people who already form groups, yet most designs still treat each person as a separate engagement target. A systematic review of 63 studies from 2000 to 2025 finds this Individual-Targeting Assumption in 60.3 percent of the corpus. Group-aware methods appear almost only after the robot is already inside an ongoing interaction; how a robot should detect, approach, and negotiate entry into a pre-formed group before first contact is left open. When individual targeting is used, three recurring failures appear: misreading inter-member signals as disengagement, skipping the group's internal agreement to engage, and neglecting bystanders. The paper therefore reframes proactive group HRI not as scaled-up one-to-one interaction but as a distinct design problem whose critical gap is the entry phase.","feed_headline":"Proactive robots treat groups as solo targets in 60% of studies","feed_subtitle":"How robots should detect and join a pre-formed group remains unstudied across two decades of work.","key_machinery":"The Individual-Targeting Assumption (ITA): the design tendency to treat each person in a social group as an independent engagement target based only on individual signals (gaze, proximity, orientation, speech) while remaining blind to the group's relational structure. It is contrasted with the Group-Aware Approach (GAA), which models at least one relational property of the group, and is mapped across interaction contexts (group entry, facilitation, spatial conduct) and three inductively coded failure modes.","core_discovery":"Across 63 proactive HRI studies in group settings, the Individual-Targeting Assumption—modeling each co-present person as an independent engagement target from individual signals alone—is present in 60.3 percent of the corpus. Group-aware approaches that model relational group properties arise almost entirely in the facilitation context after the robot is already embedded. How a robot should detect and negotiate entry into a pre-formed group before initiating contact remains unaddressed. Three failure modes—engagement misdetection, social ratification blindness, and bystander neglect—recur under individual targeting. Proactive HRI with groups is therefore not an extension of dyadic interacti","pith_inferences":["Public robots in malls, hospitals, and schools will keep failing social legitimacy until entry protocols treat groups as units that must ratify inclusion.","The same structural gap likely appears in non-robot multiparty agents that initiate contact with co-present people.","A direct test would compare robots that wait for collective orientation against robots that address the nearest individual, measuring ratification and bystander comfort.","Crowd-formation detectors already used for collision avoidance could be extended to social-openness classifiers without building perception from scratch."],"forward_implications":["Group-entry perception and negotiation become the priority target for proactive public robots.","Engagement, ratification, and bystander effects must be evaluated at the group level, not only per person.","Facilitation techniques such as turn-taking and participation balance do not transfer to entry, where the robot has only a brief observation window.","Designs that treat multiparty interaction as scaled dyadic interaction will keep producing the three documented failure modes.","Future systems need robot-centric detection of F-formations, inter-member gaze, and group openness cues during approach."],"fun_headline_variants":["60% of proactive robot studies treat groups as solo targets","Robots ignore group structure in most proactive HRI work","How robots should join pre-formed groups remains unaddressed","Individual-Targeting Assumption found in 60% of group studies","Proactive HRI misses group entry across two decades of research"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The 60.3 percent prevalence figure and the three failure counts rest on two authors' consensus inductive coding of themes without a formal reliability statistic on those theme codes, so different boundary choices could change the numbers.","fun_headline_variants_meta":{"raw":{"variants":["60% of proactive robot studies treat groups as solo targets","Robots ignore group structure in most proactive HRI work","How robots should join pre-formed groups remains unaddressed","Individual-Targeting Assumption found in 60% of group studies","Proactive HRI misses group entry across two decades of research"]},"model":"grok-4.5","effort":"low","cost_usd":0.005862,"raw_usage":{"total_tokens":1567,"prompt_tokens":837,"num_sources_used":0,"completion_tokens":87,"cost_in_usd_ticks":58620000,"prompt_tokens_details":{"text_tokens":837,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":643,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":837,"tokens_out":87,"duration_ms":7250,"temperature":1.0,"reasoning_tokens":643,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T16:42:38.811500+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Independent re-coding of the same 63 papers with the published codebook, or an expanded search that finds multiple proactive systems that explicitly model group openness and negotiate entry before first contact; a large share of group-aware entry designs would overturn the claim that entry is unaddressed.","supporting_citations":[],"review_version":1}