{"id":"4132fdca-9778-4acc-92b3-aa37e56c110c","arxiv_id":"2607.04972","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"HOLA uses hypergraphic games and open-ended partner sampling to learn multi-robot pursuit policies that adapt to unseen partners, environments, and team sizes, transferring zero-shot to physical drones and quadrupeds.","lead":"HOLA trains multi-robot teams to coordinate with unknown partners, novel maps, and changing team sizes using a hypergraph game model of who works well with whom. Real Crazyflie drones and L1 quadrupeds ran the learned policies without fine-tuning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Partner-distribution mismatch between training hypergraph and evaluation teammate pool is the load-bearing soft spot for the three-dimension claim.","rationale":"The reader correctly isolates the weakest link: the training-time partner distribution (population-internal, inverse-Myerson) is not the same as the evaluation distribution (external heterogeneous agents). That mismatch is load-bearing for the strongest claim of simultaneous adaptation across environments, partners, and scales. No other single issue (small team sizes, qualitative hardware, thin Table II margins) is as central; those are secondary limitations once partner generalization is granted. The proposed concrete test directly falsifies or supports the assumption without requiring new theory. Because the paper already shows solid multi-baseline sim gains and zero-shot hardware demos, the contribution remains accept-shaped once this gap is closed or quantified—hence CONDITIONAL is unchanged. I agree with the reader’s diagnosis and do not raise a stronger or orthogonal concern.","tokens_in":28725,"tokens_out":611,"duration_ms":5465,"concrete_test":"Retrain or re-evaluate HOLA while forcing a fraction of Oracle partners (e.g., 30–50 % of sampled π_−j) to be drawn from the exact evaluation pool (Greedy / VICSEK / D3QN-G) or frozen checkpoints of those policies, keeping all other hyperparameters fixed. Report open-team capture rate, collision rate, and AEL on both drone and quadruped platforms (Fig. 6 protocol). If success drops by more than ~10 absolute points relative to the published HOLA numbers, the partner-distribution assumption is load-bearing and the three-dimension claim weakens; if numbers hold, the concern is largely defused.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that inverse-Myerson sampling on the evolving preference hypergraph (Eqs. 12–19, ϕ Solver) plus P_team / P_env produces policies that zero-shot coordinate with the fixed heterogeneous evaluation pool (Greedy, VICSEK, D3QN-G variants; Table I) and under within-episode membership changes. Those evaluation partners are never vertices of the OH-Game used to build hyperedges or to compute ϕ; they are externally trained rule/RL agents with different action representations and skill levels. The paper therefore assumes that diversity induced inside the co-evolving population is a sufficient proxy for true open partners. If that proxy fails, the simultaneous three-dimension generalization (and the hardware transfer story that rests on it) is overstated. The ablation HOLA_R only removes the ϕ Solver while still sampling from the same population, so it does not test the distribution-mismatch hypothesis. Fixed-team environment margins are also thin (Table II: 44 % vs PBT 42.67 %), leaving little buffer if partner generalization is weaker than claimed.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper formalizes open adaptive multi-robot teaming as simultaneous generalization to unseen environments, unknown partners, and variable (including within-episode) team sizes. It introduces an Open Hypergraphic-Form Game (OH-Game / O-HyFoG) that models higher-order coalition utilities via hyperedges, derives an Open Preference Hypergraph and hyper-preference centrality η, and proposes HOLA: pre-training a diverse population (max-entropy), then iteratively expanding via a Grapher (hyperedge returns) and Oracle (approximate best-preferred agents) whose ϕ Solver samples partners by inverse Myerson/Shapley values on those returns (Eqs. 12–19). Evaluation is multi-robot cooperative pursuit on multi-drone and multi-quadruped platforms under fixed- and open-team protocols, against MAPPO, DACOOP-A, self-play, PBT, FCP, and MEP, with an ablation HOLA_R removing the ϕ Solver; policies are claimed to transfer zero-shot to Crazyflie and Zsibot L1 hardware.","tokens_in":29071,"tokens_out":1478,"duration_ms":18993,"significance":"Simultaneous three-axis open teaming is a genuine deployment bottleneck for multi-robot systems, and most prior work (ad hoc teamwork, ZSC, dynamic team size) treats axes in isolation or stays in discrete game benchmarks. A game-theoretic hypergraph formulation (explicitly not a GNN architecture) plus open-ended partner/environment expansion is a coherent design, and the multi-platform sim protocol with multi-seed error bars, heterogeneous teammate pool, and ϕ-solver ablation is stronger than typical robotics MARL papers. Direct hardware transfer without fine-tuning, if quantitatively substantiated, would be a clear contribution. Strengths to credit: external baselines and held-out partners/environments (low circularity), explicit ablation of the sampling module, and dual embodiment (aerial + legged).","major_comments":[{"comment":"Central claim vs. partner-distribution mismatch (§IV-B Oracle/ϕ Solver, Eqs. 12–19; §V-E, Table I; §VI-A/B). Training builds hyperedges and ϕ only over the co-evolving HOLA population; evaluation partners (Greedy, VICSEK, D3QN-G variants) are external rule/RL agents never present as vertices of the OH-Game. The three-dimension generalization claim therefore rests on an untested proxy assumption: that inverse-Myerson sampling inside the population induces policies that zero-shot coordinate with true open partners of different action representations and skill levels. HOLA_R only removes ϕ while still sampling from the same population, so it does not stress-test distribution mismatch. Please either (i) include evaluation partners (or behavioral clones) as held-out vertices during training, (ii) report a controlled mismatch experiment, or (iii) substantially qualify the open-partner claim an","section":null},{"comment":"Hardware evidence does not match the abstract/conclusion strength (Abstract; §V-B; Fig. 4; §VI–VII). The manuscript asserts direct transfer to Crazyflie and Zsibot L1 without fine-tuning and “robust real-world coordination in novel environments with unseen teammates,” but §VI reports quantitative capture/collision/length results only in simulation. Fig. 4 is a setup photo; no table of real-world success rate, collision rate, episode length, or number of trials under open-team conditions is given. For a robotics journal claim of this weight, add quantitative hardware metrics under the same open-team protocol (or clearly demote the claim to qualitative demonstration).","section":null},{"comment":"Environment-axis margins under simultaneous stress are thin (Table II; Fig. 5 vs. Fig. 6). In fixed-team novel environments, HOLA SUC is 44.00% vs. PBT 42.67% and SP 40.00%, with AST essentially tied with PBT. Open-team multi-drone results look stronger, but the simultaneous three-dimension claim is load-bearing and currently uneven across axes. Clarify statistical significance (e.g., paired tests over seeds/episodes), and discuss whether env gains are partly confounded by richer three-obstacle topology (as the text itself notes higher capture in “unseen” layouts).","section":null},{"comment":"Tractability of the ϕ Solver (§IV-B, Eq. 12; Proposition 4.3; Algorithm 1). Inverse Myerson is written as an average over all permutations Π(V_j), which is factorial in population size, and the Grapher enumerates subsets for each cardinality. The paper does not state population sizes used, whether exact Shapley is replaced by sampling/Monte-Carlo, or wall-clock cost per generation. Without this, reproducibility and scalability claims for “open-ended” growth are incomplete. Specify the approximation (if any), |V|, |L|, and compute budget.","section":null}],"minor_comments":[{"comment":"Notation drift: OH-Game, O-HyFoG, Open Hypergraphic-Form Game, and “preference hypergraph OPG” are used interchangeably; pick one acronym set and stick to it (esp. Def. 4.1–4.2 and §IV-B).","section":null},{"comment":"Fig. 5 caption refers to “HOLA R (marked as Our R)” while the text uses HOLA_R / HOLAR; align labels with the legend.","section":null},{"comment":"Eq. (1)–(2) joint agent-action/type spaces use power-set notation that is easy to misread; a short example of a valid element of A_C would help.","section":null},{"comment":"Within-episode membership change mechanism in evaluation (§V-C Open Team Protocol) is described at a high level (“partners join or leave”) but not operationalized (when, how many, sampling rule). Align with Algorithm 1’s P_team sampling.","section":null},{"comment":"Related work on multi-robot pursuit and ZSC is solid; a brief pointer to recent continuous-control ad hoc / open-team robotics work (beyond Hanabi/Overcooked) would better situate the hardware claim.","section":null},{"comment":"arXiv footer and journal header show placeholder dates/volume; clean for camera-ready.","section":null}],"recommendation":"major_revision","confidential_remarks":"Fit for a robotics/MARL journal is good if hardware numbers and the partner-mismatch discussion are fixed. The hypergraphic framing is more of a credit-assignment/sampling device than a deep new game solution concept; that is fine if sold accurately. I would not reject on novelty grounds alone. Watch that abstract claims do not outrun §VI evidence after revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a competent cs.RO/MARL paper that actually ships the thing people keep talking about: simultaneous open partners, open maps, and within-episode scale changes, plus zero-shot transfer to Crazyflie and L1 quadrupeds. That combination is rare enough to matter.\n\nWhat is new is the game-theoretic packaging, not a GNN. They define an open hypergraphic-form game, a preference hypergraph, hyper-preference centrality, and inverse-Myerson partner sampling inside an open-ended Grapher/Oracle loop (HOLA). The formulation is cleanly separated from message-passing architectures, and the training deliberately expands partner and environment diversity rather than fixing a team. Empirically they do the right comparisons: MAPPO, DACOOP-A, self-play, PBT, FCP, MEP; fixed vs open team; seen/unseen layouts; two morphologies; and an ablation of the ϕ solver. Effect sizes favor HOLA on capture rate and usually on collisions and episode length, with multi-seed bars in the fixed-team figure. Hardware is more than a photo: same policies, no fine-tuning, two embodiments.\n\nThe soft spot the stress-test flags is real and load-bearing. Evaluation teammates (Greedy, VICSEK, D3QN-G) are external agents that never sit as vertices in the OH-Game used to build hyperedges or ϕ. The paper assumes co-evolving population diversity is a good enough proxy for true open partners. HOLA_R only removes the ϕ solver while still sampling inside that population, so it does not test the mismatch. Fixed-team environment margins are also thin (44 % vs PBT 42.67 %). Team scales stay small and hardware remains qualitative. None of that collapses the contribution, but it means the “all three dimensions simultaneously” claim is stronger than the partner-distribution evidence.\n\nMath and citation pattern look fine for an empirical robotics paper; circularity is low. Reproducibility is hurt by missing code/data. This is for people working on ad-hoc multi-robot coordination or zero-shot multi-agent RL who care about physical transfer. I would send it to referees; it deserves a serious review, not a desk reject. I would cite the problem formalization and the dual-platform result; I would not treat the partner-generalization story as settled until someone stress-tests the distribution gap.","headline":"Solid multi-robot MARL systems paper: hypergraphic game + open-ended training, real dual-platform hardware, but partner-pool mismatch is the soft underbelly of the three-dimension claim.","tokens_in":29725,"tokens_out":602,"would_cite":true,"duration_ms":5867,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Robot teams can learn to coordinate at once with unknown partners, new environments, and team sizes that change mid-task.","keywords":["open adaptive teaming","multi-robot collaboration","cooperative pursuit","hypergraphic-form game","zero-shot coordination","multi-drone","multi-quadruped","sim-to-real transfer"],"falsifier":"Train HOLA as described, then evaluate mid-episode team changes against partners whose behaviors sit outside the training population and evaluation pool (Greedy, VICSEK, D3QN-G variants); if capture rate, collisions, and episode length collapse relative to in-pool open-team tests, the three-axis generalization claim fails.","tokens_in":29573,"feed_emoji":"🤖","tokens_out":662,"duration_ms":7769,"temperature":0.7,"pith_summary":"Real multi-robot work rarely keeps the same teammates, the same map, and the same headcount. This paper names that joint requirement open adaptive multi-robot teaming and argues that closed training with fixed partners is the wrong default. It models cooperation as a hypergraphic-form game, so payoffs attach to whole coalitions rather than only pairs, and uses that structure to decide which partners to train against as team membership shifts inside an episode. The resulting algorithm, HOLA, grows partner and environment diversity over training instead of optimizing for one fixed lineup. On cooperative pursuit with drones and quadrupeds, the method beats the listed baselines on all three axes and runs on physical Crazyflie and Zsibot platforms without fine-tuning.","feed_headline":"Robots learn to team with strangers as crew size changes","feed_subtitle":"A coalition-game trainer beats fixed-team methods on drones and quadrupeds, then runs on hardware without fine-tuning.","key_machinery":"Open hypergraphic-form game (and its preference hypergraph with hyper-preference centrality): a game-theoretic model in which hyperedges carry coalition utilities for variable team sizes; HOLA’s Oracle uses inverse Myerson-style values on that hypergraph to sample hard partners while environment and team-size distributions keep expanding.","core_discovery":"The authors claim that open adaptive multi-robot teaming—simultaneous zero-shot coordination with unseen partners, novel environments, and variable team sizes including within-episode joins and leaves—can be solved by a hypergraphic-form game that scores multi-agent coalitions, combined with open-ended training that keeps expanding partner and environment diversity. Their algorithm HOLA, built on that formulation, outperforms standard multi-agent and population-based baselines across those three dimensions on multi-drone and multi-quadruped pursuit, and the learned policies transfer directly to hardware without retuning.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Hypergraphic games let robots team with strangers of any size","HOLA trains robots for open teams that change mid-mission","Robots adapt to new partners and sizes via hypergraphic games","Open-ended training builds robot teams for unseen crewmates","Game-theoretic HOLA enables multi-robot teams that reshape on the fly"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That training against partners chosen by inverse cooperative value on the learned preference hypergraph, plus sampling of environments and team sizes, is enough to generalize to truly open partners and mid-episode membership changes that were never part of that population.","fun_headline_variants_meta":{"raw":{"variants":["Hypergraphic games let robots team with strangers of any size","HOLA trains robots for open teams that change mid-mission","Robots adapt to new partners and sizes via hypergraphic games","Open-ended training builds robot teams for unseen crewmates","Game-theoretic HOLA enables multi-robot teams that reshape on the fly"]},"model":"grok-4.5","effort":"low","cost_usd":0.004104,"raw_usage":{"total_tokens":1240,"prompt_tokens":786,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":41040000,"prompt_tokens_details":{"text_tokens":786,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":383,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":786,"tokens_out":71,"duration_ms":3115,"temperature":1.0,"reasoning_tokens":383,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T10:42:27.748205+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train HOLA as described, then evaluate mid-episode team changes against partners whose behaviors sit outside the training population and evaluation pool (Greedy, VICSEK, D3QN-G variants); if capture rate, collisions, and episode length collapse relative to in-pool open-team tests, the three-axis generalization claim fails.","supporting_citations":[],"review_version":1}