{"total":14,"items":[{"citing_arxiv_id":"2606.25526","ref_index":24,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2026-06-24T08:05:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Proposes a clipping objective for sequential trust-region updates in independent-actor cooperative MARL that yields a monotonic improvement bound and sub-linear convergence to epsilon-Nash equilibria while reducing advantage variance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25867","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"CINOC: Cardinality-Invariant Neural Operator Policies for Scalable PDE Control","primary_cat":"eess.SY","submitted_at":"2026-05-25T13:56:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"CINOC policies trained on small agent populations exhibit cardinality invariance enabling zero-shot transfer to larger populations in PDE control via mean-field theory.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.18024","ref_index":8,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2026-05-18T08:14:38+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"IBAL framework constructs information-theoretic adversarial attacks on agent observations and actions to train MARL agents that remain robust to interaction disruptions and agent-missing scenarios.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12655","ref_index":123,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning","primary_cat":"cs.AI","submitted_at":"2026-05-12T19:01:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MAVIC corrects Bellman backups at instruction boundaries by adjusting the incoming objective and restoring continuation value, enabling consistent estimation under stochastic instruction switching in cooperative MARL.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12555","ref_index":28,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DelAC: A Multi-agent Reinforcement Learning of Team-Symmetric Stochastic Games","primary_cat":"cs.MA","submitted_at":"2026-05-11T12:00:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Team-symmetric games always have team-symmetric Nash equilibria solvable via linear complementarity problems, and the DelAC actor-critic MARL algorithm outperforms existing methods in simulations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.09212","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2026-05-09T23:14:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MARS replaces additive clipping and soft penalties in multi-agent trust-region methods with a symmetric geometric barrier, matching or exceeding MAPPO and MASPO performance across 47 tasks in eight environments.","context_count":1,"top_context_role":"method","top_context_polarity":"extend","context_text":"multiplicatively symmetric penalty shape and a penalty weight αt calibrated to the desired expansion or contraction target. Combining the penalty in Equation (7) and the aligned weight in Equation (11) with the base ratio surrogate in Equation (9) gives the final per-agent MARS objective: LMARS i,t (θ) =r i,t(θ)bAt −α t \u0012 ri,t(θ) + 1 ri,t(θ) −2 \u0013 . (12) Thus, MARS keeps the standard ratio-weighted policy-gradient term but replaces additive clipping or additive quadratic penalties with a calibrated multiplicatively symmetric barrier. The full actor objective averages this surrogate over agents and timesteps in the same CTDE training loop as MAPPO and MASPO. 6 MAPPO MASPO MARS (Ours) (a) AeroJAX (b) JaxNav"},{"citing_arxiv_id":"2605.08391","ref_index":51,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SACHI: Structured Agent Coordination via Holistic Information Integration in Multi-Agent Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2026-05-08T19:00:34+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SACHI enriches agent representations via graph transformer convolutions over inter-agent graphs to enable holistic information integration, outperforming baselines across five cooperative tasks with statistical significance.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"tasks,\" inProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS), 2021. [Online]. Available: http://arxiv.org/abs/2006.07869 [50] M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar, N. Nardelli, T. G. J. Rudner, C.-M. Hung, P. H. S. Torr, J. Foerster, and S. Whiteson, \"The StarCraft Multi-Agent Challenge,\"CoRR, vol. abs/1902.04043, 2019. [51] J. Hu, S. Wang, S. Jiang, and M. Wang, \"Rethinking the implementation tricks and monotonicity constraint in cooperative multi-agent reinforcement learning,\" inThe Second Blogpost Track at ICLR 2023, 2023. [Online]. Available: https: //openreview.net/forum?id=Y8hONVbMSDj [52] R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare, \"Deep reinforcement learning at the"},{"citing_arxiv_id":"2605.03842","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SOAR: Real-Time Joint Optimization of Order Allocation and Robot Scheduling in Robotic Mobile Fulfillment Systems","primary_cat":"cs.AI","submitted_at":"2026-05-05T15:09:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SOAR is a unified DRL method using soft allocations, event-driven MDP, and heterogeneous graph transformers that cuts global makespan by 7.5% and average order completion time by 15.4% at sub-100ms latency in RMFS.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.18190","ref_index":10,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Scalable Neighborhood-Based Multi-Agent Actor-Critic","primary_cat":"cs.LG","submitted_at":"2026-04-20T12:45:59+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MADDPG-K scales centralized critics in multi-agent RL by limiting each critic to k-nearest neighbors under Euclidean distance, yielding constant input size and competitive performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2603.23964","ref_index":83,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments","primary_cat":"cs.AI","submitted_at":"2026-03-25T05:56:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"An empirical literature analysis reveals a bifurcation in RL environments into Semantic Prior (LLM-dominated) and Domain-Specific Generalization ecosystems with distinct cognitive fingerprints.","context_count":1,"top_context_role":"dataset","top_context_polarity":"background","context_text":"Big 2[79] Self-learning Card Game Agents 2018 arXiv:1808.10442 Part II: Complex Simulation, Physics & Logistics (2019-2023) Google Research Football[80] Simulated Soccer & Strategy 2019 arXiv:1907.11180 Overcooked-AI[81] Human-AI Coordination & Puzzles 2019 arXiv:1910.05789 EPyMARL[82] Grid-world Foraging 2020 arXiv:2006.07869 Robot Warehouse (RW ARE)[83] Multi-Robot Warehouse Logistics 2020 arXiv:2006.07869 Habitat 3.0[84] Interactive & Human-Robot Synergy 2023 arXiv:2310.13724 MA-Gym Cooperative Grid-world Settings 2021 GitHub: ma-gym VMAS[85] Vectorized 2D Physics Control 2022 arXiv:2207.03530 Isaac Gym[86] GPU-accelerated Physics Simulation 2021 arXiv:2108.10470 Part III: Standardized Suites & Hardware Acceleration (2020-2023)"},{"citing_arxiv_id":"2511.20857","ref_index":103,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory","primary_cat":"cs.CL","submitted_at":"2025-11-25T21:08:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Evo-Memory is a new streaming benchmark and evaluation framework for self-evolving memory in LLM agents, unifying over ten memory modules and introducing the ReMem pipeline for continual improvement on multi-turn and reasoning datasets.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2511.14135","ref_index":39,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AdaFair-MARL: Enforcing Adaptive Fairness Constraints in Multi-Agent Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2025-11-18T04:48:50+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AdaFair-MARL enforces workload fairness as an explicit second-order cone constraint in cooperative MARL via adaptive primal-dual optimization, achieving near-perfect constraint satisfaction while preserving team performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2508.01049","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies","primary_cat":"cs.LG","submitted_at":"2025-08-01T20:07:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"CoSER adaptively samples joint actions in CTDE MARL to reduce sampling error relative to the joint on-policy distribution, empirically improving reliability of independent policy gradient convergence.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2502.03506","ref_index":16,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Optimistic {\\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning","primary_cat":"cs.MA","submitted_at":"2025-02-05T12:06:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Optimistic ε-Greedy Exploration adds decoupled optimistic networks that converge in probability to maximum returns and samples from them with probability ε to increase optimal joint-action frequency in CTDE MARL.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}