{"id":"73cdcf68-3457-4be9-9ee8-d0fc10a166be","arxiv_id":"2507.10142","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey proposing adaptability as a three-part taxonomy (learning, policy, scenario-driven) for organizing and evaluating MARL under changing conditions.","lead":"This survey organizes multi-agent reinforcement learning research under a new adaptability framework with three dimensions: learning, policy, and scenario-driven adaptability. It offers researchers a common vocabulary for describing how algorithms handle shifting agent populations, changing tasks, and unfamiliar partners.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 contradicts the prose in Section 3.2, so the taxonomy's promised reliable classification is not established; without a defined scoring rubric, the central claim is conditional at best.","rationale":"The strongest claim is not that any particular algorithm is adaptable, but that the taxonomy itself is a valid organizing lens and that applying it to existing algorithms and benchmarks yields reliable classifications. For a survey whose contribution is a framework, the framework's own worked application is the evidence. Table 1 is the clearest place to check that evidence, and it fails. The inconsistencies are internal to the paper, not disagreements with outside consensus, so they directly undermine the central claim in its current form. I do not recommend rejection: the taxonomy is plausible, the coverage is broad, and the contradictions are localized to the operationalization, so a careful revision could restore the contribution. I agree only partially with the reader's weakest-assumption diagnosis: the deeper issue is less abstract separability than the absence of an operational rule for determining which dimension a shift belongs to and what rating a paradigm receives. The offline asynchronous-execution example shows the phase confusion the reader worried about, but the most damaging evidence is the specific Table 1 versus Section 3.2 mismatches. The proposed concrete test would settle whether the inconsistencies are merely typographical or symptomatic of a framework that cannot be applied reliably. The current conditional verdict is appropriate: accept only after the authors reconcile Table 1 with their own prose and state explicit classification criteria.","tokens_in":26983,"tokens_out":7544,"duration_ms":93031,"concrete_test":"Re-score Table 1 using only the suitability language in Section 3 prose, with a pre-registered mapping (e.g., 'natively suitable' implies a checkmark, 'support effective cooperation' implies at least a triangle, 'fundamentally misaligned' implies a cross), and report every cell that cannot be derived from the prose. In addition, have two independent annotators classify the same rows using only Section 3 definitions and report inter-annotator agreement. The concern lands if Networked MARL's Cooperative entry remains a cross, Centralized Critic's Competitive entry remains a checkmark, or Independent Learning's Competitive entry remains a checkmark without an explicit exception in Section 3.2.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the three-dimensional adaptability taxonomy yields a reliable, practically grounded way to classify MARL algorithms and benchmarks. The load-bearing condition is that the classification can actually be applied consistently. That condition fails already in Table 1, the paper's own operationalization, which contradicts Section 3.2. Section 3.2 says 'networked architectures also support effective cooperation in large-scale settings', but Table 1 rates Networked MARL's Cooperative column as incompatible. Section 3.2 says CTDE methods are 'fundamentally misaligned with competitive settings', yet Table 1 rates Centralized Critic's Competitive column as natively suitable. Section 3.2 says IL and mean-field methods are 'usable only in restricted forms' for competitive tasks, yet Table 1 rates Independent Learning's Competitive column as natively suitable. No rubric defines what the symbols mean or how prose statements map to ratings. The boundary between dimensions is also phase-confused: Table 1 is presented under Learning Adaptability, but Offline MARL's Asynchronous Execution rating is justified by deployment-time inference behavior. Because the framework's own application is internally inconsistent, the review does not yet deliver what the central claim promises; it delivers a plausible taxonomy whose classifications need repair and explicit criteria.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes adaptability as an assumption-aware lens for evaluating MARL algorithms under shifting conditions, organized into three dimensions: learning adaptability (training/paradigm robustness), policy adaptability (reuse and generalization of a single policy), and scenario-driven adaptability (benchmark and evaluation design). It reviews classical paradigms (CTDE, independent learning, offline, mean-field, networked, model-based, safe MARL), five families of policy-generalization methods, and a broad set of structured-game, simulator, and LLM-based benchmarks. Two summary tables operationalize the framework: Table 1 rates learning adaptability of paradigms across seven axes, and Table 2 characterizes environment configurability. The central claim is that this taxonomy is a valid, practically grounded way to organize MARL research and to obtain reliable classifications of algorithms and benchmarks.","tokens_in":27224,"tokens_out":7120,"duration_ms":79347,"significance":"If the inconsistencies are repaired, the framework is a useful conceptual contribution: it gives a shared vocabulary for concepts such as scalability, robustness, generalization, and transferability that are often used inconsistently, it separates learning-time from deployment-time questions in a principled way, and it collates a wide literature with specific pointers to algorithms and benchmarks. The paper's explicit key questions and the benchmark table (Table 2) are valuable resources. The main value is organizational rather than novel technical results, which is appropriate for a survey. The weakness is that the framework's own operationalization in Table 1 is not yet reliable, so the central claim is conditional on a revision.","major_comments":[{"comment":"The table and prose give contradictory classifications. Table 1 rates Independent Learning's Competitive entry as natively suitable, while Section 3.2 says IL is 'usable only in restricted forms and typically lack mechanisms for anticipating adversarial strategies'; Centralized Critic's Competitive entry is natively suitable, while Section 3.2 says CTDE methods are 'fundamentally misaligned with competitive settings'; and Networked MARL's Cooperative entry is incompatible, while Section 3.2 says networked architectures 'support effective cooperation in large-scale settings.' Since the paper's central claim is that the three-dimensional taxonomy yields reliable classifications, these direct contradictions are load-bearing and must be resolved by aligning the table with the prose or by revising the prose.","section":"Table 1 vs. Section 3.2"},{"comment":"The sentence 'As shown in Table 1, IL and networked MARL offer the greatest flexibility for distributed, asynchronous deployment' is not what Table 1 shows: Networked MARL has Distributed Training = ✓ but Asynchronous Execution = ✗. The summary overstates the table and should either be corrected or the table's Async column for Networked MARL should be changed to a partial rating with a justification in the text.","section":"Section 3.3 final paragraph"},{"comment":"The rating for Offline MARL (△) is justified in Section 3.3 by deployment-time inference behavior, but Table 1 is presented as a dimension of Learning Adaptability. This mixes training-phase and deployment-phase criteria within a single column, which is exactly the kind of phase confusion the framework's learning/policy distinction was designed to avoid. Each row or column should be tagged with the phase to which the claim applies, or the column should be split into training-time and deployment-time asynchrony.","section":"Table 1, Async. Exec. column"},{"comment":"No scoring rubric is given for the four-level suitability scale. In particular, IL's Cooperative = ✓ sits uneasily with Section 3.1's statement that IL methods 'typically struggle to learn coordinated behaviours in tightly coupled environments.' The authors should specify what evidence (benchmark results, architectural properties, or both) upgrades a paradigm to 'natively suitable' and how border cases map to 'partially suitable'; without such a rubric the table is not falsifiable.","section":"Section 3.2 and Table 1"}],"minor_comments":[{"comment":"Section 3.3 contains the incomplete sentence 'These constraints' immediately before 'Distributed Training Support.'","section":"Section 3.3"},{"comment":"The phrase 'we introduce a structured framework describe MARL adaptability' should read 'a structured framework that describes MARL adaptability.'","section":"Section 2.2"},{"comment":"The sentence 'Competitive tasks involvedirectlyconflictinggoals' should be 'Competitive tasks involve directly conflicting goals.'","section":"Section 3.2"},{"comment":"The sentence 'Another approaches propose an auto-curriculum' should be 'Another approach proposes an auto-curriculum' or 'Other approaches propose an auto-curriculum.'","section":"Section 4.4"},{"comment":"The sentence 'CTDE methods, such as VD, CC, and HA, are only partially compatible' with distributed training is inconsistent with Table 1, which gives Value Decomposition Distributed Training = ✗; either the text or the table should be adjusted.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"No circularity concern: the self-citations are descriptive and do not drive the framework. The central problem is the internal consistency of Table 1 with the surrounding prose; I view this as fixable by adding an explicit scoring rubric and reconciling table and text. If the authors cannot produce a consistent Table 1, the survey's practical value as a reference would be significantly reduced. The benchmark table (Table 2) is a useful resource and should be retained."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my quick read.\n\nThe genuinely new bit is the three-dimensional taxonomy: learning adaptability, policy adaptability, and scenario-driven adaptability. It separates what changes, when it changes, what adaptation is allowed, and what success means, and it does a decent job of positioning existing terms like robustness, generalization, and transferability relative to each other. The survey also covers a wide range of recent work, including offline pretraining and zero-shot coordination, and the benchmark table in Section 5 is a handy resource.\n\nThe soft spot is exactly where the stress-test lands: Table 1 does not match the prose. Section 3.2 says networked architectures 'also support effective cooperation in large-scale settings,' but Table 1 marks Networked MARL as incompatible with cooperative tasks. The same section says CTDE methods are 'fundamentally misaligned with competitive settings,' yet Table 1 gives Centralized Critic a native-suitable check for competitive. IL and mean-field are described as 'usable only in restricted forms' competitively, while Table 1 rates IL as natively suitable. There is also no stated rubric for what the symbols mean, so the ratings are unfalsifiable as they stand. The phase confusion is real too: Table 1 sits under learning adaptability, but the Offline MARL asynchronous-execution rating is justified by deployment-time inference.\n\nNone of this kills the taxonomy itself. The conceptual framework is plausible and could give the community a common vocabulary. But the survey's central promise is that applying the framework yields reliable classifications, and that promise fails at the first operational step. The fix is straightforward: define explicit criteria for the three ratings, apply them consistently, and reconcile the table with the prose. That probably requires a revision rather than acceptance as is.\n\nWho is this for? MARL researchers thinking about evaluation and benchmark design. They will get value from the taxonomy and the bibliography even if they cannot trust Table 1 in its current form. I would send it to review, not desk reject, because the contribution is real and the problems are repairable. My own verdict: revise and resubmit, with the rubric as a required addition.","headline":"Useful taxonomy, unreliable operationalization: Table 1 contradicts its own prose, but the framework deserves a serious referee.","tokens_in":27725,"tokens_out":4170,"would_cite":true,"duration_ms":41733,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes adaptability as the unifying lens for judging whether multi-agent reinforcement learning algorithms survive shifting conditions, and divides it into learning adaptability, policy adaptability, and scenario-driven…","keywords":["multi-agent reinforcement learning","adaptability","learning adaptability","policy adaptability","scenario-driven adaptability","zero-shot coordination","offline MARL","benchmark evaluation"],"falsifier":"Take a single MARL algorithm, hold its architecture fixed, and construct two environments that differ only in whether a population change happens during training or during deployment. If the algorithm's failure mode is identical in both, with the same coordination breakdown and the same performance drop curve, then the division between learning and policy adaptability is not separating what it claims to separate, and the three dimensions are not independent.","tokens_in":26799,"feed_emoji":"🔄","tokens_out":3726,"duration_ms":38927,"temperature":0.7,"pith_summary":"This survey argues that the reason multi-agent reinforcement learning (MARL) works in benchmarks but fails in deployments is that algorithms carry unstated assumptions about agent count, training access, synchronization, and partner behavior. The paper introduces adaptability as a single umbrella concept for what changes when a MARL system moves to a new setting, and splits it into three dimensions: learning adaptability (training-time resilience to population, task, and execution changes), policy adaptability (reuse of a trained policy across tasks, roles, and unfamiliar partners), and scenario-driven adaptability (whether benchmarks expose those shifts in a controlled way). The claim is that organizing existing work under these three headings clarifies how scalability, robustness, generalization, and transferability relate, and reveals where evaluation is underspecified. A sympathetic reader would take the paper's contribution to be the framework itself: a vocabulary and checklist for asking which assumption a MARL algorithm violates and what kind of adaptation is allowed.","feed_headline":"One framework splits MARL reliability into three adaptability axes","feed_subtitle":"Learning, policy, and scenario shifts are separated so benchmarks and algorithms can be compared honestly.","key_machinery":"The central object is the three-dimensional adaptability taxonomy itself: learning adaptability, policy adaptability, and scenario-driven adaptability, each subdivided into axes such as population scaling, task structure, execution constraints, permutation invariance, offline-to-online transfer, and zero-shot coordination. Its work is to act as a classification scheme: every MARL paradigm and benchmark is assigned a position on these axes, and the paper's tables turn implicit assumptions into explicit, comparable entries. The framework also supplies four clarifying questions—what changes, when the change occurs, what adaptation is allowed, and what success means—which the paper uses to separate concepts that previous reviews conflate.","core_discovery":"The paper's central claim is that adaptability, defined as any change in environment dynamics during learning or execution, can serve as a unified and practically grounded lens for evaluating MARL reliability, and that this lens decomposes into three separable dimensions. Learning adaptability asks whether the learning paradigm itself remains stable when agent populations, task structures, or execution constraints shift. Policy adaptability asks whether a policy trained on one configuration can generalize or transfer to new tasks, roles, or partners without retraining. Scenario-driven adaptability asks whether benchmarks and evaluation protocols expose controlled, diagnostic shifts. The paper then applies this taxonomy across paradigms, across policy mechanisms, and across benchmarks, producing tables that classify each approach's suitability along the three axes. The intended contribution is unification: existing notions such as scalability, robustness, and transferability become facets of adaptability, and their intersections mark where real-world readiness is actually tested.","pith_inferences":["The taxonomy's three dimensions likely interact more than the paper's presentation suggests; for example, asynchronous execution can change effective population scale, and offline data can encode partner conventions, so a single algorithm may be classifiable on two axes simultaneously.","A testable extension is an adaptability 'scorecard' that forces each study to state the violated assumption, the dimension, the allowed adaptation budget, and the success metric, a card that could be extracted directly from the tables proposed here.","The framework could be applied retroactively to single-agent RL, where policy adaptability and scenario-driven adaptability are already studied but learning adaptability under population scaling disappears, suggesting adaptability is really a continuum rather than MARL-specific.","Building a benchmark suite that varies each of the three dimensions independently while measuring performance degradation would empirically validate or refute the separability assumption that the taxonomy rests on."],"forward_implications":["If the taxonomy is adopted, evaluation of MARL algorithms will routinely report which adaptability dimension is being tested, rather than relying on a single aggregate score.","Existing properties like scalability, robustness, and transferability become partial facets of adaptability; a method that is scalable but not policy-adaptable is located on the framework rather than praised or dismissed wholesale.","Benchmark designers gain concrete design principles, such as incremental population scaling, role diversification, reward consistency, and held-out partner diversity, for making environments diagnostically useful.","Offline-to-online transfer and zero-shot coordination are identified as under-specified axes where both algorithms and benchmarks lag, which points future work toward those gaps.","The classification predicts that no single existing paradigm covers all three dimensions natively; independent learning is most flexible at execution, centralized methods at coordination, and neither at cross-task generalization."],"supporting_citations":[{"why":"Supplies the theory-and-algorithms baseline of MARL that the adaptability framework reorganizes.","marker":"[8]"},{"why":"Represents the prior survey on cooperative MARL in open environments whose fragmented terms the framework aims to unify.","marker":"[2]"},{"why":"Defines the scalability facet for large-population MARL that adaptability subsumes.","marker":"[7]"},{"why":"QMIX serves as the representative centrally coupled CTDE method used to motivate learning adaptability limits.","marker":"[14]"},{"why":"MADDPG and the MPE benchmark anchor the analysis of task structures and population scaling.","marker":"[33]"},{"why":"Mean-field MARL is the canonical abstraction-based paradigm illustrating the scalability-coordination trade-off.","marker":"[13]"},{"why":"Networked MARL grounds the discussion of distributed and asynchronous execution constraints.","marker":"[24]"},{"why":"Offline MARL (ICQ-MA) anchors the offline pretraining and offline-to-online policy adaptability discussion.","marker":"[21]"},{"why":"Other-Play is the foundational zero-shot coordination method that defines the policy adaptability facet for unfamiliar partners.","marker":"[125]"},{"why":"Overcooked provides the benchmark evidence for zero-shot coordination with multiple viable conventions.","marker":"[129]"}],"fun_headline_variants":["MARL reliability split into three adaptability axes","Adaptability taxonomy: learning, policy, and scenario shifts","Three-axis framework clarifies MARL benchmark honesty","New review: MARL evaluation needs assumption-aware axes","Separate learning, policy, scenario shifts for MARL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that changes in learning conditions, deployment conditions, and evaluation scenarios can be cleanly separated into three independent dimensions; if real shifts straddle or transform one another, as when asynchronous execution reshapes population scale or offline data encodes partner conventions, the taxonomy will misclassify or double-count the same phenomenon.","fun_headline_variants_meta":{"raw":{"variants":["MARL reliability split into three adaptability axes","Adaptability taxonomy: learning, policy, and scenario shifts","Three-axis framework clarifies MARL benchmark honesty","New review: MARL evaluation needs assumption-aware axes","Separate learning, policy, scenario shifts for MARL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1300,"prompt_tokens":923,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":303}},"tokens_in":539,"tokens_out":377,"duration_ms":4532,"temperature":1.0,"reasoning_tokens":303,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:38:01.566328+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single MARL algorithm, hold its architecture fixed, and construct two environments that differ only in whether a population change happens during training or during deployment. If the algorithm's failure mode is identical in both, with the same coordination breakdown and the same performance drop curve, then the division between learning and policy adaptability is not separating what it claims to separate, and the three dimensions are not independent.","supporting_citations":[],"review_version":1}