{"id":"6c9588ba-76f1-435f-963b-4bbc94322676","arxiv_id":"2504.21048","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that categorizes recent MARL algorithms for resource allocation optimization by application domain, training paradigm, and challenge type, without introducing new experimental results.","lead":"This paper reviews how multi-agent reinforcement learning is used to allocate resources such as bandwidth, energy, and computing power across telecommunications, energy grids, transportation, and manufacturing. It organizes recent work into a taxonomy by application domain and training paradigm, and lists publicly available simulation benchmarks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"This survey's 'comprehensive review' claim is unauditable: Section 4.1 never documents search databases, keywords, inclusion criteria, or a time window, and Table 3 diverges from the section's own prose in several rows.","rationale":"I read this as a survey whose central claim is comprehensiveness: 'This survey provides a comprehensive review of recent MARL algorithms for RAO... a structured taxonomy' (abstract), reinforced in Section 1 by the claim that existing surveys miss this intersection and that 'This survey fills that gap.' For the claim to hold, the papers summarized in Section 4.1 and encoded in Table 3 must be representative of the MARL-for-RAO literature, and the taxonomy must be a faithful compression of that corpus. The reader's weakest_assumption identifies the first condition, and I agree it is the most load-bearing point. My read sharpens it with evidence: Section 4.1 never states how papers were located, 'recent' and 'high-impact' are undefined, and Table 3 diverges from the survey's own prose in several rows (Renewable Energy, Smart Grid, Mobile Edge Computing, Autonomous Vehicles). Also, Jain and Kumar (2023) is described once as fully centralized DRL and once as MAAC. None of this falsifies the survey, but it shows the corpus and taxonomy are not auditable, which is precisely what a 'comprehensive... structured taxonomy' must be. I looked for stronger objections and did not find one. The foundations (MDP, POMDP, Dec-POMDP, Bellman equations) are standard and correctly stated. The four benchmarks have working public URLs consistent with the cited sources, and ContainerGym is honestly disclosed as an RL environment that can be extended to MARL. Performance statements such as 'SAC outperforms other baselines' (Section 4.2.1) are attributed to the cited works rather than asserted as new findings. The concern is therefore about methodology and representativeness (a correctness risk for the taxonomy's prevalence and challenge assignments), not about agreement with field consensus. The concrete test, reconstructing the corpus with a stated query and measuring citation-ranked coverage, would settle the concern in days. Because the fix is cheap and non-destructive, I recommend keeping the reader's CONDITIONAL verdict unchanged: add a methodology paragraph or soften the comprehensiveness claim. This is an agreement with the reader's weakest_assumption, not a new objection.","tokens_in":31766,"tokens_out":13672,"duration_ms":118955,"concrete_test":"Reconstruct the Section 4.1 corpus with an explicit protocol and measure coverage: query Scopus or Web of Science for 2019-2024 with TITLE-ABS-KEY(('multi-agent reinforcement learning') AND ('resource allocation' OR 'task offloading' OR 'resource management')), restrict to the survey's five application domains, and rank by citations. Take the top 30 papers per domain (150 total) and compute the fraction present in the survey's reference list. If more than roughly 25% of the domain top-30 are absent, the corpus is not representative and the 'comprehensive' claim should be withdrawn or qualified. Within the same pass, audit Table 3 row-by-row against the survey's own Section 4.1 prose; the Renewable Energy (MAAC), Smart Grid (MADQN), and Mobile Edge Computing (MATD3, Com-DDPG) rows already fail this audit and should be corrected regardless of the coverage outcome.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 1 claim this survey fills a gap by providing a comprehensive review of recent MARL for RAO, with Table 3 packaging the corpus into a structured taxonomy by application, algorithm, and challenge. The load-bearing condition for that claim is that the Section 4.1 corpus is representative of the MARL-for-RAO literature. The paper never establishes this: there is no documented search strategy, no database list, no keyword set, no inclusion or exclusion criteria, and no time window defining 'recent.' The Introduction's phrase 'selecting high-impact studies' is not operationalized. Because the taxonomy rows and prevalence statements (e.g., Section 4.2.3's 'VD has been rarely used in RAO') are derived from this undocumented corpus, a differently selected corpus could plausibly yield a different taxonomy, so the comprehensiveness claim is not auditable. The concern is corroborated by internal inconsistencies: Table 3 is not derivable from the text it summarizes. Renewable Energy prose includes MAAC (Jayanetti et al. 2024) but the table omits it; Smart Grid prose includes MADQN (Kumari et al. 2024) but the table lists only MAPPO; Mobile Edge Computing prose includes MATD3 (Zhao et al. 2022) and Com-DDPG (Gao et al. 2023) but the table omits both; the Autonomous Vehicles row lists MADDPG and MADQRL with no textual support. Jain and Kumar (2023) is also described as fully centralized DRL in Section 4.2.1 and as MAAC in Section 4.1.3. These mismatches mean the taxonomy is not tightly auditable even from the survey's own corpus. This is a method gap, not a detected falsity. The four benchmark subsections give working public URLs consistent with the cited source papers, and ContainerGym (Section 5.4) is honestly disclosed as single-agent RL that can be extended to MARL. The MDP, Bellman, and Dec-POMDP foundations are standard. The fix is feasible: add a survey-methodology subsection or soften 'comprehensive' to 'a representative sample.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews recent multi-agent reinforcement learning (MARL) methods for resource allocation optimization (RAO). The paper introduces classical RAO approaches and their limitations, gives a brief formal foundation of RL and MARL (MDPs, stochastic games, Dec-POMDPs, training paradigms), and then surveys MARL applications across telecommunications, energy, distributed computing, transportation, and manufacturing. It organizes the literature into an application-based taxonomy (Table 3), associates the surveyed works with four RAO challenge categories (adaptability, partial observability, large scale, heterogeneity), lists four publicly available RAO-related benchmarks with code links, and closes with future research directions.","tokens_in":32099,"tokens_out":3161,"duration_ms":29279,"significance":"If its corpus is representative, the survey fills a genuine gap by consolidating recent MARL-for-RAO work into a structured taxonomy and pointing practitioners to four usable benchmarks (BSK-RL, Power Distribution Networks, CityFlow, ContainerGym). The standard RL/MARL equations in Section 3 are reproduced correctly, and the domain summaries generally match the stated contributions of the cited papers. The benchmark section is a practical strength because it ships code links and names tested algorithms in Table 4. However, the survey's central claim of being 'comprehensive' is currently not auditable: the paper-selection process is undocumented, and the main taxonomy table is not derivable from the prose that accompanies it. These issues weaken the paper as a reference until the methodology and internal consistency are fixed.","major_comments":[{"comment":"The 'comprehensive review' claim is not auditable because the paper never documents its literature-selection process. Section 4.1 presents the corpus in Table 3 without stating the search databases, keyword sets, inclusion/exclusion criteria, or the time window covered by 'recent'; Section 1's phrase 'selecting high-impact studies' is not operationalized. Since the taxonomy and prevalence statements in Section 4.2 are derived entirely from this undocumented corpus, a differently selected corpus could plausibly produce a different taxonomy. The authors should add a methodology subsection describing the search and selection protocol.","section":"Section 4.1 / Section 1"},{"comment":"Table 3 is not derivable from the prose it summarizes. For example, the Renewable Energy row omits MAAC (Jayanetti et al. 2024) even though Section 4.1.2 explicitly describes it; the Smart Grid row lists only MAPPO while Section 4.1.2 discusses MADQN (Kumari et al. 2024); the Mobile Edge Computing row omits MATD3 (Zhao et al. 2022) and Com-DDPG (Gao et al. 2023), both described in Section 4.1.3; and the Autonomous Vehicles row lists MADDPG and MADQRL with no supporting discussion in Section 4.1.4. The reader cannot trust the taxonomy until the table and text are reconciled.","section":"Table 3 vs Section 4.1.2"},{"comment":"There is an internal contradiction about Jain and Kumar (2023). Section 4.2.1 describes the work as a fully centralized MARL method evaluated with DQN, DDPG, and SAC, while Section 4.1.3 lists it as using the MAAC algorithm. These are incompatible descriptions of the same paper, and this inconsistency undermines the reliability of both the challenge analysis and the application taxonomy.","section":"Section 4.2.1 vs Section 4.1.3"},{"comment":"The statement 'Currently, VD has been rarely used in RAO' is an unsupported prevalence claim. It is based on a single citation (Ahmed et al. 2023) and there is no quantitative count of value-decomposition papers in the corpus, which is itself undocumented. Either qualify the statement as an observation about the present selection or provide a systematic count.","section":"Section 4.2.3"}],"minor_comments":[{"comment":"The displayed constraint 'Pn i=1 xi ≤ N' appears with unrendered LaTeX markup; please format it properly.","section":"Section 2.1.1"},{"comment":"Several citations repeat the author name awkwardly, e.g., 'Chen et al. Chen et al. (2022)', 'Wang et al. Wang et al. (2021)', and 'Xu et al. Xu et al. (2020)'; these should be cleaned to a single citation form.","section":"Section 4.1.2"},{"comment":"Halabian (2019a) and Halabian (2019b) share the same title and DOI but are listed as two distinct entries; please merge or differentiate them.","section":"References"},{"comment":"The 'Traffic Management' entry in Table 4 lists 'Graph based MARL' without naming a specific algorithm; since the surrounding text mentions CityFlow and PressLight, please specify which algorithm was actually tested.","section":"Section 5.3 / Table 4"}],"recommendation":"major_revision","confidential_remarks":"This is a useful survey with a practical benchmark section, but the 'comprehensive review' claim currently rests on an undocumented corpus, and Table 3 has verifiable mismatches with the prose. I would ask for a transparent selection methodology and a full reconciliation of Table 3 with Section 4.1 before considering acceptance. The issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this survey is worth refereeing, but only after the authors make their literature selection auditable and fix Table 3. The MARL foundations are correctly reproduced, the benchmark pointers are the most useful part, and the resource classification in Section 2 is a reasonable organizing device. The 'comprehensive review' claim, however, rests on a corpus selection that is never documented.\n\nWhat is genuinely useful: Section 3 gives a clean, standard account of MDPs, Dec-POMDPs, and the CTCE/DTDE/CTDE paradigms. Section 5 lists four public benchmarks with working URLs and honestly notes that ContainerGym is single-agent RL that can be extended to MARL. For a practitioner looking for entry points into MARL-for-RAO, the application-by-application organization helps.\n\nThe soft spots are real but fixable. Section 4.1 never states search databases, keywords, inclusion criteria, or a time window. The introduction says 'selecting high-impact studies' but does not operationalize that. Consequently, the taxonomy and prevalence claims (e.g., 'VD has been rarely used in RAO') are not auditable. More concretely, Table 3 diverges from the section's own prose in several rows: Renewable Energy prose includes MAAC (Jayanetti 2024) but the table omits it; Smart Grid prose includes MADQN (Kumari 2024) but the table lists only MAPPO; Mobile Edge Computing prose includes MATD3 (Zhao 2022) and Com-DDPG (Gao 2023) but the table omits both; the Autonomous Vehicles row lists MADDPG and MADQRL with no textual support. Also, Jain and Kumar (2023) is described as fully centralized DRL in Section 4.2.1 and as MAAC in Section 4.1.3. These inconsistencies matter because the table is the paper's main synthesis. The performance statements like 'outperforming static or rule-based methods' are borrowed from the cited papers without critical assessment; that is common in surveys, but the authors should present them as reported claims, not established facts.\n\nRecommendation: send to peer review with a request for a methodology subsection or a softened comprehensiveness claim, plus a corrected Table 3. The topic deserves a survey, and this is a decent draft rather than a finished reference. With those fixes it could be a genuinely useful entry point for practitioners and newcomers.","headline":"A useful but uneven survey: the benchmark pointers and MARL foundations are solid, but the 'comprehensive' claim is not yet auditable because the literature selection is undocumented and Table 3 disagrees with the prose.","tokens_in":32708,"tokens_out":2689,"would_cite":false,"duration_ms":24668,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of MARL for resource allocation builds a taxonomy by challenge and training paradigm.","keywords":["multi-agent reinforcement learning","resource allocation optimization","survey","taxonomy","centralized training decentralized execution","partial observability","scalability","benchmarks"],"falsifier":"Run a systematic, preregistered literature search on MARL for resource allocation with explicit inclusion criteria and a fixed time window; if the resulting corpus contains major application areas or algorithm families that the survey's four-challenge taxonomy cannot accommodate, the survey's representativeness claim is false.","tokens_in":31580,"feed_emoji":"🤖","tokens_out":8470,"duration_ms":78982,"temperature":0.7,"pith_summary":"Multi-agent reinforcement learning (MARL), the paper argues, is the right tool for resource allocation optimization (RAO) in modern dynamic systems because it targets the four weaknesses of classical methods: rapid change, partial observability, large scale, and heterogeneity. The survey organizes recent work across telecommunications, energy, distributed computing, transportation, and manufacturing into a structured taxonomy based on training paradigms (CTCE, DTDE, CTDE) and recurring challenges. It identifies centralized training with decentralized execution (CTDE) as the most flexible and widely used paradigm for RAO, and it lists four publicly available benchmarks for evaluating algorithms before deployment. If this mapping is correct, the survey gives practitioners a practical starting point for choosing a MARL approach and gives researchers a shared reference for where the gaps are.","feed_headline":"Survey maps the MARL toolkit for dynamic resource allocation","feed_subtitle":"Three training paradigms and four recurring challenges give practitioners a starting point for choosing an algorithm.","key_machinery":"The load-bearing organizing device is the three-way taxonomy of MARL training and execution paradigms: centralized training and centralized execution (CTCE), decentralized training and decentralized execution (DTDE), and centralized training with decentralized execution (CTDE). The survey uses this taxonomy, together with the Decentralized Partially Observable Markov Decision Process (Dec-POMDP) formalism, to map each application to a mechanism: CTCE for small fully coordinated systems, DTDE for privacy-sensitive or highly distributed settings, and CTDE as the workhorse that combines global coordination with local autonomy. The other half of the machinery is the four-category challenge schema, which groups the literature by adaptability, partial observability, scalability, and heterogeneity.","core_discovery":"The paper's central claim is that MARL is the appropriate toolkit for modern RAO because it answers exactly the four limitations that sink classical methods: rapid change, partial observability, large scale, and heterogeneity. The survey organizes recent literature into a taxonomy by application domain and by training paradigm, and it identifies a set of four public benchmarks. It also reports that CTDE is the dominant paradigm for RAO, since centralized training gives agents a global view during learning while decentralized execution keeps them scalable and responsive at deployment time. The paper presents this as a synthesis of the current research landscape rather than as a new algorithm or experimental result.","pith_inferences":["A testable design rule follows from the survey's examples even though the paper does not state it: choose DTDE when communication cost or privacy dominates, and CTDE when system-level optimality dominates.","Because the four benchmarks use different tasks and metrics, the survey cannot support cross-domain algorithm rankings; a unified RAO benchmark suite would be the next step toward validating its taxonomy.","The recurring appearance of graph-based MARL across domains suggests that relational structure, not raw state vectors, may be the common substrate of RAO, and a cross-domain transfer experiment could test this directly.","The survey's challenge categories imply that partial observability and adaptability are coupled: partial information slows adaptation, so methods that improve observability may improve adaptability without extra machinery."],"forward_implications":["CTDE emerges as the default starting paradigm for RAO problems where both global coordination and scalability matter, because it trains with global information and executes on local observations.","The four public benchmarks give the field a common testbed for comparing MARL algorithms before deployment, covering satellite tasking, power-grid voltage control, traffic signal control, and container-based waste processing.","Classical methods such as linear programming, heuristics, and game theory should be treated as baselines rather than competitors in dynamic RAO, since the survey identifies their core limitations as static assumptions, centralization, and poor scaling.","Graph-based MARL appears across energy, manufacturing, and mobile networks, indicating that modeling inter-agent relationships is becoming a standard tool for RAO.","Value decomposition methods such as QMIX are available for credit assignment in cooperative RAO but remain underused compared with centralized-critic methods, and the paper notes there is no guarantee they converge to a global optimum."],"supporting_citations":[{"why":"Supplies the RL/MDP formalization, including value functions and Bellman equations, on which the survey's MARL foundation is built.","marker":"Sutton and Barto (2018)"},{"why":"Provides the Dec-POMDP formulation that the survey uses to model partial observability in decentralized RAO.","marker":"Oliehoek et al. (2016)"},{"why":"Introduces MADDPG, the centralized-critic CTDE algorithm that the survey presents as a core approach for RAO.","marker":"Lowe et al. (2017)"},{"why":"Introduces QMIX value decomposition, the credit-assignment method the survey discusses for cooperative RAO.","marker":"Rashid et al. (2020)"},{"why":"Demonstrates MAPPO's effectiveness in cooperative MARL, which the survey cites across several RAO applications.","marker":"Yu et al. (2022)"},{"why":"Provides a review of MARL challenges and solutions that the survey draws on to frame non-stationarity and partial observability.","marker":"Nguyen et al. (2020)"},{"why":"Supplies the active-voltage-control power network benchmark used for evaluating MARL algorithms in RAO.","marker":"Wang et al. (2021)"},{"why":"Supplies the city-scale traffic signal control environment used as a large-scale RAO benchmark.","marker":"Zhang et al. (2019)"},{"why":"Supplies the waste-processing resource-allocation benchmark, one of the four public testbeds the survey lists.","marker":"Pendyala et al. (2024)"},{"why":"Supplies the satellite tasking environment used as a spacecraft-oriented RAO benchmark.","marker":"Stephenson and Schaub (2024b)"}],"fun_headline_variants":["MARL survey flags four pitfalls classical RAO hits","Taxonomy of MARL for resource allocation: survey of paradigms","Why MARL wins at resource allocation: survey of paradigms","CTDE dominates MARL resource allocation, survey finds","Four challenges that make MARL the RAO toolkit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim to be comprehensive depends on the papers it selected for Section 4.1 being representative of the field, and it does not state a search strategy, inclusion criteria, or time window.","fun_headline_variants_meta":{"raw":{"variants":["MARL survey flags four pitfalls classical RAO hits","Taxonomy of MARL for resource allocation: survey of paradigms","Why MARL wins at resource allocation: survey of paradigms","CTDE dominates MARL resource allocation, survey finds","Four challenges that make MARL the RAO toolkit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001111,"raw_usage":{"total_tokens":4540,"prompt_tokens":766,"completion_tokens":3774,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":382,"completion_tokens_details":{"reasoning_tokens":3695}},"tokens_in":382,"tokens_out":3774,"duration_ms":28971,"temperature":1.0,"reasoning_tokens":3695,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:31:31.038288+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic, preregistered literature search on MARL for resource allocation with explicit inclusion criteria and a fixed time window; if the resulting corpus contains major application areas or algorithm families that the survey's four-challenge taxonomy cannot accommodate, the survey's representativeness claim is false.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the RL/MDP formalization, including value functions and Bellman equations, on which the survey's MARL foundation is built."},{"cited_title":"Samvelyan, C.S","cited_arxiv_id":null,"evidence_quote":"Introduces QMIX value decomposition, the credit-assignment method the survey discusses for cooperative RAO."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the active-voltage-control power network benchmark used for evaluating MARL algorithms in RAO."},{"cited_title":"Dettmer, T","cited_arxiv_id":null,"evidence_quote":"Supplies the waste-processing resource-allocation benchmark, one of the four public testbeds the survey lists."}],"review_version":1}