{"id":"08ab119b-5928-480d-8123-655036d112de","arxiv_id":"2607.09993","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"PR-MRE hedges regret over high-confidence type subsets and, via PRMRE-PSRO, yields more distribution-shift-robust team policies than BNE on graph Capture-the-Flag.","lead":"The paper defines Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE) for adversarial team games with hidden types, then learns approximate strategies via a robust double-oracle PSRO method with an SDP meta-solver. It shows on graph Capture-the-Flag that the resulting policies scout before committing and hold up better under type-distribution shifts than Bayesian Nash policies.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"SDP relaxation fidelity is unvalidated for the Graph CtF metagame that carries the empirical claim","rationale":"The reader already flags the SDP fidelity and the thin experimental base as the weakest assumptions. The single most load-bearing technical gap is the unvalidated quality of the SDP relaxation on the very metagame that produces the headline empirical figures: without a gap or rank report, one cannot know whether the Blue policies in Figs. 8–9 are responding to an approximate PR-MRE or to an arbitrary mixture. The threat-model modeling choice is secondary; even under a perfect threat model the certificate does not transfer if the meta-solver is inaccurate. The proposed test is cheap (tiny normal-form instance) and decisive. No stronger internal inconsistency appears; the theory under exact solution is sound. Hence the verdict remains CONDITIONAL, with the same emphasis the reader already placed on strengthening the SDP validation and multi-seed experiments.","tokens_in":28314,"tokens_out":513,"duration_ms":5074,"concrete_test":"On the exact 2\times4 seed payoff tensor of Experiment 4.1, solve the non-convex PRMRE-EQM-NF bilinear program (or a high-order Moment-SOS hierarchy) to global optimality; compare the resulting Blue mixture and worst-case regret value against the McCormick SDP + rank-1 projection used in the paper. If the regret gap exceeds ~5–10 % of the type-optimal range or the support of the mixture differs materially, the empirical robustness claim is no longer supported by the theory.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Lemma 4 and the tightest-lower-bound claim hold for exact PR-MRE. The empirical claim that PRMRE-PSRO yields substantially better worst-case win-rates (Figs. 6–9) rests on the McCormick-cut SDP + rank-1 projection (A.2) producing a meta-equilibrium whose regret vector is close enough to the true robust bilinear optimum that the subsequent MAPPO BR inherits the certificate. The paper never reports duality gap, recovered rank, or post-projection residual on the actual 2\times4 Graph CtF seed metagame (only on the tiny synthetic 1\times5\times3 / 2×5×3 matrices). If the relaxation is loose, the mixture fed to Blue BR is not a PR-MRE mixture and the robustness curves cannot be attributed to the claimed concept.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE) for adversarial team games with asymmetric information. Under a typicality-preserving threat model (Definition 3.2), PR-MRE minimizes worst-case expected regret over admissible type distributions S(μ̄,δ) rather than expected payoff or fully distribution-free worst-case regret. Lemmas 1–4 establish uniform performance lower bounds; for normal-form games the equilibrium is cast as a robust bilinear program (PRMRE-EQM-NF) that is lifted to a McCormick-cut SDP with rank-1 projection (Appendix A.2). The SDP is used as a meta-solver inside a robust double-oracle / PSRO loop (PRMRE-PSRO, Algorithms 1–2) whose best responses are obtained by MAPPO on GNN policies. On a 2v2 Graph Capture-the-Flag instance the method is claimed to produce Blue policies with better worst-case win-rate under type-distribution shifts than BNE-PSRO, together with scouting-like visitation patterns.","tokens_in":28591,"tokens_out":1140,"duration_ms":8814,"significance":"If the claims hold, the work supplies a clean intermediate point on the risk-neutral / distributionally-robust / fully worst-case spectrum for asymmetric-information team games, together with an explicit programming formulation and a population-based learning algorithm. The normal-form bilinear program, the SDP lift, and the lower-bound certificates (Lemmas 1–4) are mathematically self-contained and correctly derived from standard regret algebra and water-filling. The Graph CtF experiments illustrate a concrete behavioral consequence (scouting before commitment) that is of practical interest for multi-agent path-finding and reachability games. These contributions are novel relative to existing ATG and PSRO literature and would be of interest to the algorithmic game-theory and multi-agent RL communities, provided the empirical link between the SDP meta-solver and the claimed robustness certificate is tightened.","major_comments":[{"comment":"The central empirical claim (Figs. 6–9, §4.1) attributes improved worst-case win-rates under distribution shift to PR-MRE. That attribution rests on the McCormick-cut SDP + rank-1 projection (Appendix A.2) producing a meta-equilibrium whose regret vector is close to the true robust bilinear optimum of PRMRE-EQM-NF. The paper never reports duality gap, recovered rank, residual after projection, or any other fidelity metric on the actual 2×4 Graph CtF seed metagame that carries the claim; only the tiny synthetic 1×5×3 / 2×5×3 matrices are solved. Without such diagnostics it is impossible to know whether the mixture fed to the MAPPO best-response is a genuine PR-MRE mixture or an uncontrolled approximation, and therefore whether the robustness curves can be credited to the proposed concept.","section":null},{"comment":"Empirical support is limited to a single expansion iteration on a hand-crafted 2×4 seed metagame (Figs. 6–7) with no multi-seed statistics, no error bars, and no comparison against a pure MRE (δ=0) baseline or against DRO-MaxMin on the same structured ambiguity set. The performance-robustness curve (Fig. 8) and the visitation densities (Fig. 9) are therefore suggestive but insufficient to establish that PRMRE-PSRO systematically discovers more robust strategies than risk-neutral PSRO on graph-structured ATGs.","section":null},{"comment":"The typicality-preserving threat model (Definition 3.2) is presented as the appropriate model of strategic deception, yet the paper supplies no external justification or sensitivity analysis for the choice of δ, nor does it examine how the recovered Blue mixture changes when the low-probability covers are misspecified. Because Lemma 4’s “tightest lower bound” certificate holds only inside S(μ̄,δ), the practical value of the certificate depends on whether that set is a realistic description of the adversary’s power; this assumption remains untested outside the synthetic matrices.","section":null}],"minor_comments":[{"comment":"Notation for the low-probability cover set oscillates between Θℓ, Θl and L; a single consistent symbol would improve readability.","section":null},{"comment":"Figure 3’s caption refers to “δ3” and “δ4” which appear to be typographical errors for θ3 and θ4.","section":null},{"comment":"The relationship between the BR_MMR stopping criterion used in Algorithm 2 and the exact PR-MRE best-response operator (Definition 3.3) is only sketched; a short formal statement would clarify that the training loop is still approximating the same fixed point.","section":null},{"comment":"Several references to “Appendix A.1” for proofs of Lemmas 1 and 4 are correct, but the main text could briefly indicate that the proofs are elementary rearrangements so that a reader need not consult the appendix for the core argument.","section":null}],"recommendation":"major_revision","confidential_remarks":"The mathematical core (threat model, bilinear program, SDP lift, lower-bound lemmas) is sound and publishable. The empirical section is the clear weak point; if the authors can add even modest SDP-fidelity diagnostics on the Graph CtF metagame and a multi-seed robustness comparison, the paper would be a solid contribution. Without those additions I would be reluctant to accept for a top venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real addition is a previously empty cell: regret-based best response under coarse trust in the nominal, via typicality-preserving covers S(μ̄,δ). That is not just renaming MRE or DRO. They give a robust bilinear program, a McCormick SDP lift, and a double-oracle/PSRO meta-solver that can actually be run with MAPPO best responses. The taxonomy tables and the 1×5×3 / 2×5×3 matrices make the concept differences concrete; Lemmas 1–4 are standard regret algebra and are correctly stated under the threat model they chose.\n\nWhat works: the lower-bound certificate (Lemma 4) is the right kind of guarantee for the setting they care about, and the Graph CtF behavioral story (scouting both corridors instead of committing to the majority flag) is easy to read and matches the objective. The normal-form math and the SDP construction look solid on the page.\n\nSoft spots, in proportion: the empirical claim is thin. One expansion iteration on a hand-crafted 2×4 seed metagame, no multi-seed stats, no code. More importantly, the stress-test lands: they never report duality gap, recovered rank, or post-projection residual on the actual Graph CtF metagame that carries Figures 6–9—only on the tiny synthetic matrices. So the robustness curves cannot yet be tightly attributed to exact PR-MRE rather than to “some robust mixture from a relaxed meta-solver.” The threat model itself is a modeling choice, not validated as the right model of strategic deception. They flag the open EFG/sequential question themselves.\n\nThis is for people working on robust MARL, adversarial team games, and search/security settings who need something between BNE and pure worst-case. It is serious work, not a rehash. I would send it to peer review; a good referee will demand multi-seed runs and SDP diagnostics on the metagame that supports the main claim. Worth engaging if that is your area.","headline":"Clean middle-ground equilibrium (PR-MRE) and a usable PSRO meta-solver; theory holds under its threat model, but the Graph CtF claim rests on one seed iteration and an unvalidated SDP relaxation.","tokens_in":29175,"tokens_out":527,"would_cite":true,"duration_ms":10703,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"PR-MRE is a regret-based equilibrium that protects the uninformed team against strategic shifts in hidden types while still using a nominal prior, and it can be learned via a robust double-oracle loop.","keywords":["adversarial team games","asymmetric information","minimax regret","PR-MRE","PSRO","distribution shift","graph capture-the-flag","robust equilibrium"],"falsifier":"On the same Graph CtF seed metagame, run identical PSRO expansions with BNE (or pure MRE) as meta-solver instead of PR-MRE; if Blue win-rate for test mass on the majority flag ≤ 0.6 no longer exceeds the BNE curve, or if the symmetric corridor-scouting visitation pattern disappears, the central empirical claim fails.","tokens_in":29188,"feed_emoji":"🏁","tokens_out":1024,"duration_ms":17048,"temperature":0.7,"pith_summary":"In adversarial team games one side knows a hidden type (for example a flag location) and can condition play on it; the other side only has a prior. Bayesian Nash strategies that trust that prior can be badly exploited when the type distribution is shifted. This paper introduces Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE): the uninformed team minimizes worst-case expected exploitability only over distributions that keep rare types rare under the nominal prior. That yields a tighter performance lower bound than fully distribution-free minimax regret, while still covering adversarial redistribution of mass among typical types. For normal-form games the equilibrium is a robust bilinear program with a tractable semidefinite relaxation; the authors embed that relaxation as the meta-solver in a double-oracle algorithm (PRMRE-PSRO) that expands populations with deep RL best responses. On graph Capture-the-Flag the resulting Blue policies scout both corridors instead of over-committing to the majority flag and keep higher win-rates when the type distribution moves.","feed_headline":"PR-MRE hedges hidden types better than Bayesian Nash","feed_subtitle":"On graph Capture-the-Flag, learned policies scout both corridors and keep win-rate when the prior shifts.","key_machinery":"Probabilistically Robust Minimax-Regret Equilibrium (PR-MRE): the Blue operator that minimizes the supremum of expected exploitability over the typicality-preserving threat set S(nominal, δ). For finite normal-form games this is a robust bilinear program; an SDP lift with McCormick cuts and rank-1 projection is used as the meta-solver inside a robust double-oracle / PSRO loop with RL best responses.","core_discovery":"PR-MRE is the fixed point of a Blue best response that minimizes worst-case expected exploitability over the typicality-preserving set of type distributions together with the informed Red best response. Among the concepts compared, it supplies the tightest uniform lower bound on expected payoff for every distribution inside that set, and the SDP-relaxed double-oracle PRMRE-PSRO produces Blue policies whose worst-case win-rate under type shifts substantially exceeds that of Bayesian-Nash PSRO on the Graph CtF instance.","pith_inferences":["The same typicality-preserving threat model could be applied to other hidden-parameter multi-agent settings (unknown skill tiers, map layouts) without requiring divergence-ball ambiguity sets.","Whether the McCormick SDP relaxation stays tight enough for larger type spaces will decide how far PRMRE-PSRO scales beyond two-flag Graph CtF.","A sequential (extensive-form) refinement of PR-MRE that admits local regret decomposition would open a path to CFR-style algorithms for the same robustness objective."],"forward_implications":["Uninformed-team strategies hedge across high-confidence type subsets rather than specializing to the modal type under the nominal prior.","Learned policies exhibit scouting before commitment, lowering exploitability when the realized type is not the majority hypothesis.","Performance lower bounds hold for any redistribution of mass that keeps rare types rare (low-probability covers stay low-probability).","The same meta-solver can replace the usual Nash meta-solver in PSRO for any normal-form abstraction of an asymmetric-information team game.","Fully distribution-free minimax-regret equilibrium is recovered as the special case δ = 0."],"fun_headline_variants":["PR-MRE tightens worst-case win-rate over BNE under type shifts","Probabilistically robust minimax-regret hedges hidden types better than BNE","PRMRE-PSRO yields CtF policies with higher worst-case payoff across types","Minimax-regret equilibrium resists deception in asymmetric team games","Beyond Bayesian Nash: PR-MRE bounds regret over high-confidence type sets"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The argument needs strategic deception to only redistribute mass while keeping types that are rare under the nominal prior still rare, and needs the SDP relaxation of the robust bilinear program to be accurate enough that the learned policies inherit that certificate.","fun_headline_variants_meta":{"raw":{"variants":["PR-MRE tightens worst-case win-rate over BNE under type shifts","Probabilistically robust minimax-regret hedges hidden types better than BNE","PRMRE-PSRO yields CtF policies with higher worst-case payoff across types","Minimax-regret equilibrium resists deception in asymmetric team games","Beyond Bayesian Nash: PR-MRE bounds regret over high-confidence type sets"]},"model":"grok-4.5","effort":"low","cost_usd":0.003418,"raw_usage":{"total_tokens":1231,"prompt_tokens":896,"num_sources_used":0,"completion_tokens":104,"cost_in_usd_ticks":34180000,"prompt_tokens_details":{"text_tokens":896,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":231,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":896,"tokens_out":104,"duration_ms":2666,"temperature":1.0,"reasoning_tokens":231,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T01:10:36.491048+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same Graph CtF seed metagame, run identical PSRO expansions with BNE (or pure MRE) as meta-solver instead of PR-MRE; if Blue win-rate for test mass on the majority flag ≤ 0.6 no longer exceeds the BNE curve, or if the symmetric corridor-scouting visitation pattern disappears, the central empirical claim fails.","supporting_citations":[],"review_version":1}