{"id":"c5eaeb93-8ca8-4866-b64d-8da1e9ed1277","arxiv_id":"2507.05469","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The first MOASEI competition evaluated four submitted multi-agent policies on wildfire and cybersecurity tracks, with GNN and CNN approaches matching a simple baseline.","lead":"The MOASEI 2025 competition benchmarked multi-agent AI policies in three open-world environments where agents and tasks appear, disappear, and change over time. This technical report describes the first run of the competition, the four submitted solution methods, and what the results suggest about adapting to openness.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5's openness-robustness claims rest on three configurations per domain, with no quantified openness metric; the asserted WS1-WS3 openness gradient is never operationalized.","rationale":"The paper is a competition report, and its descriptive claims—that the competition was run and four solutions were evaluated—are credible. The authors are appropriately cautious in noting that the top submissions did not statistically significantly beat the simple 'smallest' baseline. However, the paper's broader contribution, stated in the abstract and Section 5, is that the results provide empirical insight into generalization and adaptation in open environments. That contribution depends on the three held-out configurations being representative of the openness dimension the competition is designed to test. The paper asserts that WS3 has 'significantly more openness' but supplies no metric, no sampling procedure, and no coverage analysis. With only three configurations, any observed performance difference across them could be confounded by unrelated parameter changes. This is the load-bearing soft spot: if the configurations do not actually vary openness in a controlled way, the Section 5 findings do not follow. The reader's weakest_assumption identifies the same issue, and I agree. The paper should be accepted only on condition that this evaluation evidence is provided or the claims are weakened accordingly.","tokens_in":10434,"tokens_out":7418,"duration_ms":75095,"concrete_test":"Define quantitative openness metrics for each domain (e.g., rate of agent appearance/disappearance, task arrival/departure rate, distribution of agent lifetimes). Generate at least 10 held-out configurations per domain that vary these openness metrics independently while holding other environment parameters fixed, and re-run the full evaluation of all submitted and baseline policies. If the reported performance trends across WS1-WS3/CS1-CS3 do not correlate with the openness metric, or if the relative ranking of policies changes across the sampled configurations, then the Section 5 robustness claims are not supported and should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The evaluation protocol in Section 4 fixes three held-out configurations per domain and uses them to support Section 4.1's statement that WS3 has 'significantly more openness' and Section 5's claims that GNN and CNN policies are robust to agent dropout and dynamic tasks. No openness metric is defined or reported, so the WS1-WS3 performance differences cannot be attributed to openness rather than to changes in other environment parameters (e.g., fire spread rate, world size, agent count). The paper also does not describe how the three configurations were sampled from the environment family, nor does it report any variance across configuration draws. Because the competition's stated purpose is to evaluate behavior under openness, the claim that the results reveal strategies for 'generalization and adaptation in open environments' is underdetermined by the evidence provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical report documents the organization and results of the inaugural MOASEI Competition at AAMAS 2025, a multi-agent benchmarking event built on the free-range-zoo environment suite. Three tracks were designed—Wildfire, Rideshare, and Cybersecurity—with formalizations as partially observable stochastic games. Eleven teams registered and four submitted solutions: three for Wildfire (an LLM-based entry, a GNN-based entry, and a CNN-based entry) and one for Cybersecurity (a weighted-scoring entry). Each policy was evaluated with 256 runs per configuration on three held-out configurations per domain, and results were compared against naïve baselines using Wilcoxon signed-rank tests. The report's main empirical observations are that the top Wildfire submissions performed similarly to each other and to the simple \"smallest\" baseline, that the Cybersecurity winner's advantage over heuristic baselines was not statistically significant, and that no submissions were received for Rideshare. The paper also presents qualitative findings about GNN, CNN, LLM, and predictor-augmented approaches and outlines future competition plans.","tokens_in":10680,"tokens_out":4849,"duration_ms":57971,"significance":"If the descriptive claims are taken at face value, the paper documents a useful, reproducible open-agent benchmarking infrastructure and transparently reports participation and outcomes. The public codebase, fixed-seed evaluation pipeline, and candid reporting of crashes and non-significant comparisons are strengths. However, the central empirical insight is limited by the paper's own admission that the top Wildfire and Cybersecurity submissions were statistically indistinguishable from simple greedy baselines. Consequently, the stronger claims in the abstract and Section 5 about architecture-specific robustness to openness and 'promising strategies' are not established by the reported evidence. The paper's value as a competition infrastructure report is credible, but its empirical contribution requires substantial reframing or additional analysis.","major_comments":[{"comment":"The paper's own statistical comparison undercuts its central empirical claim. In Table 1, markov_mayhem (11.36±0.61) and university_of_tehran (11.41±0.63) have total Wildfire rewards essentially identical to the 'smallest' baseline (11.35±0.57), and the text in Section 4.1 states that these submissions 'did not have statistically significant differences between each other and smallest.' Section 5's claims that GNN-based policies 'demonstrated high adaptability to agent dropout and dynamic tasks' and that CNN policies 'matched the peak performance' therefore go beyond the evidence; the submissions matched, rather than outperformed, a trivial greedy heuristic. The authors should either provide additional analyses (e.g., task-specific metrics, temporal dynamics, or direct comparisons not confounded by configuration difficulty) or substantially temper the conclusions in the abstract, Section 4.1, and Section 5.","section":"Section 4.1, Table 1; Section 5"},{"comment":"The robustness-to-openness claims rest on an unquantified and unvalidated openness gradient. Section 4.1 attributes the WS1-to-WS3 differences to 'significantly more openness,' and Section 5 draws conclusions about 'robustness against openness' for GNN and CNN methods. However, the paper defines no openness metric, reports no values for agent/task entry and exit rates or other openness parameters, and does not describe how the three held-out configurations were sampled from the environment family. The observed performance differences could equally be due to changes in fire spread rate, world size, agent count, or other environment parameters. The authors should operationalize openness, report the concrete parameter settings for WS1–WS3 and CS1–CS3, or restrict all such claims to 'performance on the three tested configurations.'","section":"Section 4, first paragraph; Section 4.1; Section 5"},{"comment":"The statistical methodology for the significance claims is underspecified. The text states that n=256 independent simulation runs per policy were executed with fixed seeds, and that Wilcoxon signed-rank comparisons were used. A signed-rank test requires paired observations, so the pairing structure must be clarified: if the 256 runs are paired across policies by shared seed, describing them as 'independent' is misleading; if they are not paired, the Wilcoxon signed-rank test is inappropriate. Additionally, no multiple-comparison correction is reported for the many pairwise tests across baselines, configurations, and metrics. Because the paper's only comparative statements are negative (no statistically significant differences), the validity of the test procedure is load-bearing and should be fixed or explicitly justified.","section":"Section 4, first paragraph; Figures 3–5 and 13–15"}],"minor_comments":[{"comment":"The manuscript contains multiple typos and grammatical errors, including 'questions questions' (Section 3.2), 'signficance' (Section 4), 'recieved' and 'signficant' (Section 4.2), 'thourougly' (Section 7), and 'these modification will allow use' (Section 6). A careful proofreading pass is needed.","section":"Section 3.2, Section 4, Section 4.2, Section 7"},{"comment":"The bit_student row reports 'crash' for WS3, but the Total column (7.71) appears to be the sum of WS1 and WS2 only. The table should state explicitly that the total excludes WS3, and the text should explain how crashed runs were handled in the leaderboard.","section":"Table 1"},{"comment":"The reward specification 'rewards them by 2agents_required' is mathematically ambiguous; it should be written as 2 × agents_required or an equivalent explicit formula.","section":"Section 2.2"},{"comment":"The figures labeled 'Wilcoxon-signed-rank comparisons' are never described in the surrounding text—readers are not told what is being compared (e.g., which pairs of policies, whether the display shows test statistics or p-values). Adding a sentence or a detailed caption would make the figures interpretable.","section":"Section 4.1 and 4.2, Figures 3–5, 13–15"},{"comment":"The cybersecurity action-proportion tables contain rendering artifacts such as '𝑎𝑚𝑜𝑣𝑒', '𝑎𝑛𝑜𝑜𝑝', '𝑎𝑝𝑎𝑡𝑐ℎ', and '𝑎𝑚𝑜𝑛𝑖𝑡𝑜𝑟' in table headers. These should be replaced with plain-text labels such as 'move', 'noop', 'patch', and 'monitor' for readability.","section":"Tables 5–8"},{"comment":"The sentence 'We observe that neither number of burnouts nor simulation length are correlated with cumulative episodic rewards' reports visual correlations from Figures 6–8 without any coefficient or test. Either provide a quantitative correlation statistic or soften the claim to an observation from the plotted data.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about non-significance and infrastructure, but the abstract and Section 5 overstate the empirical findings. For a competition technical report, the descriptive and infrastructure content is within scope; however, the authors need to either strengthen the evidence for robustness claims or reframe the paper as a benchmarking-infrastructure report with deliberately modest empirical conclusions. The statistical-testing ambiguity should also be resolved before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a competition technical report, not a research paper, and it reads that way. The new content is the empirical benchmark data from the inaugural MOASEI competition: leaderboard scores, action distributions, and the honest observation that the best GNN and CNN wildfire policies were not statistically different from the 'fight the smallest fire' baseline. That last point is the most interesting thing in the report, and I give the authors credit for printing it plainly. They also deserve credit for building a reproducible pipeline on top of free-range-zoo, publishing the code, and comparing submissions against sensible baselines. The descriptive claims—the competition ran, four teams submitted, the environment behaved as documented—are credible. As infrastructure for a community, this is useful.\n\nThe soft spots are real but not disqualifying. The biggest one is exactly what your stress test flags: Section 5 claims robustness to openness, yet the paper never defines or measures openness. The three held-out configurations per domain are treated as an openness gradient, but nothing shows the gradient is openness rather than fire spread rate, world size, or agent count. The phrase 'WS3 has significantly more openness' appears without any operationalization. That is a load-bearing gap for the paper's stated purpose, and it is fixable by reporting an openness metric over configurations. Second, the bit_student crash in WS3 is reported as 'crash' in a table and never discussed; the reader cannot tell if it was a bug, a timeout, or an environment incompatibility. Minor but annoying. Third, participation is small (three wildfire, one cybersecurity, zero rideshare), so the 'large-scale empirical evaluation' language in Section 5 is overstatement. Fourth, the paper names winners even though the top entries were statistically tied with the baselines; it files this under 'findings' but the framing still leans toward 'success,' which undercuts the honest numbers.\n\nWho is this for? People building open-agent benchmarks and researchers using free-range-zoo. It will not change minds about GNNs or LLMs, and the non-significance result is the real takeaway. It deserves a serious referee, but the referee should push for a quantified openness analysis and a full accounting of the crash before the claims about generalization are allowed to stand.\n\nFor peer review: send it out, with revisions. The infrastructure is reusable and the non-significant result is worth publishing as a caution to the community.","headline":"A credible competition report whose empirical claims about openness outrun the evidence; the headline result is that learned policies did not beat a trivial baseline.","tokens_in":11080,"tokens_out":1263,"would_cite":false,"duration_ms":17765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The inaugural MOASEI competition benchmarked multi-agent policies under agent and task openness, finding the best learned policies statistically tied with a simple greedy baseline and an LLM-based submission crashing in the hardest…","keywords":["open agent systems","multiagent systems","benchmark competition","partially observable stochastic games","agent openness","task openness","free-range-zoo","graph neural networks"],"falsifier":"Re-running the evaluation on configurations sampled from a wider distribution of openness parameters, rather than the three fixed held-out scenarios, and observing that the statistical ties between learned and greedy policies disappear would refute the paper's generalization claims.","tokens_in":10261,"feed_emoji":"🔥","tokens_out":5607,"duration_ms":59763,"temperature":0.7,"pith_summary":"This paper reports the design, execution, and results of the first MOASEI competition, a multi-agent benchmarking event built on the free-range-zoo environment suite. The competition aimed to measure how AI policies cope with openness: agents and tasks that appear, disappear, or change behavior over time. The central empirical claim is that the leading submissions in the wildfire track, a graph neural network and a convolutional network, achieved the highest cumulative rewards but were not statistically distinguishable from a simple policy that always fights the smallest fire. In the cybersecurity track, the single submission outperformed the baselines but also did not statistically separate from heuristic policies. The paper argues these results are still informative, identifying the robustness of GNN and CNN policies and the viability of LLM-driven meta-optimization as directions for future work.","feed_headline":"GNN, CNN agents match a greedy baseline","feed_subtitle":"MOASEI 2025 benchmarked open-world policies; learned agents tied the strongest heuristic baseline.","key_machinery":"Partially observable stochastic games (POSGs) formalize each track, and the free-range-zoo environment suite provides three domain implementations—wildfire, cybersecurity, and rideshare—with mechanisms for agent openness and task openness. The evaluation pipeline runs 256 fixed-seed episodes per policy on three held-out configurations and applies Wilcoxon signed-rank tests to determine statistical significance. This machinery carries the argument by giving openness a concrete, per-domain definition and by making the comparison between learned and heuristic policies statistically explicit.","core_discovery":"The paper's central discovery is that a formal competition evaluating open-agent systems is feasible and yields reproducible evidence about policy classes: learned relational and convolutional policies match, but do not statistically exceed, a greedy baseline in the wildfire domain, and the LLM-augmented policy crashes in the most open configuration. In the cybersecurity domain, the winning weighted-scoring policy also failed to achieve a statistically significant advantage over heuristic baselines. The authors interpret this as evidence that open-world benchmarking can identify which architectural families generalize under openness, while also showing that the current testbed does not yet separate sophisticated learning from simple heuristics.","pith_inferences":["If the three held-out configurations are representative of the openness spectrum, the statistical ties suggest the benchmark tasks may need larger-magnitude or higher-frequency openness to separate learning-based policies from heuristics.","A concrete testable extension would combine the LLM meta-optimizer with a GNN base policy and compare against both components alone on out-of-distribution configurations.","The paper's claims about GNN and CNN robustness would be strengthened or undercut by rerunning the evaluation on environment families outside free-range-zoo with structurally different openness.","Reporting variance across draws of the evaluation configurations, rather than only across the 256 seeded episodes, would provide a more direct measure of generalization under openness."],"forward_implications":["The competition infrastructure—public environments, fixed-seed evaluation, and leaderboards—can be reused as a shared benchmark for open-agent systems.","Because the top learned policies statistically tied the greedy baseline, the current wildfire and cybersecurity testbeds do not yet reward sophisticated learning as strongly as intended.","The comparable performance of CNN and GNN policies in the wildfire track suggests lightweight spatial architectures are a viable alternative to relational ones in these domains.","The LLM-based meta-optimizer, despite not winning, demonstrates a route to policy adaptation that does not require end-to-end reinforcement learning.","With no rideshare submissions, the track design as it stood failed to attract participation, informing a needed redesign for future iterations."],"supporting_citations":[{"why":"Provides the free-range-zoo codebase with the three domain environments used as the competition testbeds.","marker":"[1]"},{"why":"The competition website that defined the tracks, rules, and evaluation guidelines for participants.","marker":"[2]"},{"why":"Supplies the formal definition of open agent systems and the POSG framework that structures each track.","marker":"[3]"}],"fun_headline_variants":["MOASEI 2025: Learned agents tie greedy baselines","Open-world benchmark: no statistical edge for AI","GNNs, CNNs match greedy baseline in open domains","LLM-driven policy crashes in most open MOASEI track","Heuristics still competitive in open-agent benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation treats three held-out configurations drawn from the same environment family as a sufficient measure of generalization, without analyzing how well they cover the space of possible openness levels.","fun_headline_variants_meta":{"raw":{"variants":["MOASEI 2025: Learned agents tie greedy baselines","Open-world benchmark: no statistical edge for AI","GNNs, CNNs match greedy baseline in open domains","LLM-driven policy crashes in most open MOASEI track","Heuristics still competitive in open-agent benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1280,"prompt_tokens":832,"completion_tokens":448,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":368}},"tokens_in":448,"tokens_out":448,"duration_ms":5544,"temperature":1.0,"reasoning_tokens":368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:24:27.173114+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the evaluation on configurations sampled from a wider distribution of openness parameters, rather than the three fixed held-out scenarios, and observing that the statistical ties between learned and greedy policies disappear would refute the paper's generalization claims.","supporting_citations":[{"cited_title":"free-range-zoo GitHub Repository","cited_arxiv_id":null,"evidence_quote":"Provides the free-range-zoo codebase with the three domain environments used as the competition testbeds."},{"cited_title":"Methods for Open Agent Systems Evaluation Initiative (MOASEI) 2025 Competition","cited_arxiv_id":null,"evidence_quote":"The competition website that defined the tracks, rules, and evaluation guidelines for participants."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the formal definition of open agent systems and the POSG framework that structures each track."}],"review_version":1}