Pith. sign in

Paper Citation Record · LEDGER

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage

As of 6 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.07548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07548 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T11:08:51.426575Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact18
  • verified fuzzy38
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7fa6a324-0853-4ec1-9ecc-abf1cf55bcfd · outbound

This paper cites A survey on multi-agent reinforcement learning and its application.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A survey on multi-agent reinforcement learning and its application

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.176673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:d66b6b9e070dcb596083d7674b9dd464437aecedfd9c88ef176f756db2b17e1e

Observation a6a66adf-e6fe-4cc6-8731-f95e7d3f33da · outbound

This paper cites A comprehensive survey on multi-agent reinforcement learning for connected and automated vehicles.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A comprehensive survey on multi-agent reinforcement learning for connected and automated vehicles

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.174462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:f371e301cd47ea72b83a1dafb16b58da19d87c8875747e04b28098081eadf6a1

Observation 4264b629-d50d-49ae-bf20-04ab19aea8ee · outbound

This paper cites Marlens: Understanding multi-agent reinforcement learning for traffic signal control via visual analytics.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Marlens: Understanding multi-agent reinforcement learning for traffic signal control via visual analytics

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.373302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:b38b6a255872a835cb405d0769a02853da25621d6538d62fd1438049821603bc

Observation b1ef45a9-d453-4972-8192-06249d9c3230 · outbound

This paper cites Multi-Agent Reinforcement Learning for Autonomous Driving: A Survey.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Multi-Agent Reinforcement Learning for Autonomous Driving: A Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.552234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:277a253162c23a0f91a6aaad82a7848bd6bafd9e81107ac7f18326497bf0e4ba

Observation e60134e6-dd69-48da-804b-a7c2cf4701b5 · outbound

This paper cites A survey of multi-agent deep reinforcement learning with communication.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A survey of multi-agent deep reinforcement learning with communication

Reference 6

Resolution
verified exact
doi, observed 2026-05-19T11:12:15.384062Z

Source-reported events for the cited work

correction dated 2024-03-16. Source: crossref record 10.1007/s10458-024-09644-x->10.1007/s10458-023-09633-6:correction, observed 2026-07-11T03:07:50.406185+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:fe614eedc5f22bc874f042f9cea2b5eee2a80d6089757e8a1d854e87ebfa3613

Observation e817814d-60e2-457c-9269-35bdf3f97a70 · outbound

This paper cites Deep reinforcement learning: A survey.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Deep reinforcement learning: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.159161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:602273b0734f5bad8b45a2acbf78fae99630279c043a4e3fd0f010f22d81c45b

Observation e875f140-d7b1-4094-b95b-503a2f59ff12 · outbound

This paper cites Weighted qmix: expanding monotonic value function factorisation for deep multi-agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Weighted qmix: expanding monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.161429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:256917d413796ecd655754bc08b4ed551c5dc2badb0589f3940f2c73fd31699a

Observation a9961338-6965-436c-a4eb-f5df2bba6ca5 · outbound

This paper cites Episodic multi-agent reinforcement learning with curiosity- driven exploration.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Episodic multi-agent reinforcement learning with curiosity- driven exploration

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.154801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:eddc8ff64643741a791275ede27702cf759853d4a3f73e3e3a121469f2af0960

Observation 79164b48-ec60-4a69-9a84-e5c8ae8cc421 · outbound

This paper cites Emu: Efficient episodic memory utilization of cooperative multi-agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Emu: Efficient episodic memory utilization of cooperative multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.152821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:0989032d29a7d47637b584882b01b7c0b594d6f0913b48a699867b6a89d8ca7c

Observation 1c64f3b7-67a7-4d46-be07-f370a00e785b · outbound

This paper cites A survey of reinforcement learning algorithms for dynamically varying environments.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A survey of reinforcement learning algorithms for dynamically varying environments

Reference 12

Resolution
verified exact
doi, observed 2026-05-19T11:12:15.381754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:d130f2bf67d13ad23b1c8ec610a96f939481f13ac4c39a34c40b4fbd3fabb2bb

Observation ba868521-54a0-4eeb-8d17-1215d06026e8 · outbound

This paper cites A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:12:15.533776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:fb3a331983052d72539174c68cf889317182b6077c2aa14bee94bf6fbd6354a8

Observation ecc5c207-bda1-405e-96f9-fc50c19aaaeb · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage The StarCraft Multi-Agent Challenge

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:12:15.527732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:d4053908ae65a663a910a6c44cd6d65243143150f0c41696cc8ad93e95995a27

Observation 62209474-2270-4fa6-b4cf-cfd2558de696 · outbound

This paper cites SMACv2: An improved benchmark for cooperative multi-agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage SMACv2: An improved benchmark for cooperative multi-agent reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.172558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:54094c0d9d1668d009bdb360b515fc6c01afcc9f6541ab2238a36b88f23113ce

Observation d81b91bc-55f9-4533-af4b-2bb589c21cc7 · outbound

This paper cites A survey on curriculum learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A survey on curriculum learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.144601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:85be9e01168d6ea894ae2af2c9b9e0d6ec88e941e66af6a90ea41e4fa10ecd0b

Observation 588abd56-c028-45c3-b2a8-cef14516287d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:12:15.536641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:a6756354112abb0c1246253bf3efcc8e69273c7974a5b6643e926debc1a12f59

Observation cc91dc0d-3ddb-4cec-b056-da5754061c71 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Counterfactual multi-agent policy gradients

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.140405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:479d1e7a2f6b0458c427752c3d3ebdea061b74f41c80886fbf61861ee4afc030

Observation f6b5ba46-8fad-4c5d-a37e-9bc7e13145fa · outbound

This paper cites Monotonic value function factorisation for deep multi- agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Monotonic value function factorisation for deep multi- agent reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.150809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:05b2799ec0fe25f301b217ab2ee039ffc413d62bac1b5940a90d0edbb1cff371

Observation 8457f65e-2143-4303-9cc2-fa357808a64c · outbound

This paper cites Discriminative experience replay for efficient multi- agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Discriminative experience replay for efficient multi- agent reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.142626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:27633650b31c7bd6d1fbc5b905003d14df469c15ba26fa0a108ba72df6436e49

Observation b941258a-95ae-4172-a86d-9435420f5286 · outbound

This paper cites Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Dealing with Non-Stationarity in Multi-Agent Deep Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:12:15.549194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:8bcd8c952e0ae6c514628ea2f754d1684d46839ddeed5a1d8b698e44bbb7bb5e

Observation 4278ca5b-09f7-47d7-bb87-c9db5b38ab1f · outbound

This paper cites Dealing with non- stationarity in MARL via trust-region decomposition.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Dealing with non- stationarity in MARL via trust-region decomposition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.099054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:442bde22bcc2ad8f0d6d1eb5d3fd73bceeda70e87fb419fb8033557a5d6868ac

Observation 1ca85564-b158-442b-b36b-e271b90f0eb0 · outbound

This paper cites Tackling non-stationarity in decentralized multi-agent reinforcement learning with prudent q- learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Tackling non-stationarity in decentralized multi-agent reinforcement learning with prudent q- learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.106247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:fb1a4398ac2e467355408d4066d074480f71fc6b87622a97643a0564126f7760

Observation 05aa0ca2-b337-4907-9b98-f1ac5d0df478 · outbound

This paper cites Dealing with non-stationarity in decentralized cooperative multi-agent deep reinforcement learning via multi-timescale learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Dealing with non-stationarity in decentralized cooperative multi-agent deep reinforcement learning via multi-timescale learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.166324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:f5316e102e2df05a3f241b96f4199d8b2f3355ab10214e0dd565c22de36ed8c5

Observation e709aba6-7510-4e62-9d4a-2df09df715b5 · outbound

This paper cites Monotonic improvement guarantees under non- stationarity for decentralized PPO.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Monotonic improvement guarantees under non- stationarity for decentralized PPO

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.133809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:205ba145b671ec840ef7755df8a904e7c02e07bfd29f6602cc4a2839a2e733a9

Observation e644fa80-0a46-4153-8ad2-698557c5c9f3 · outbound

This paper cites Value-decomposition networks for cooperative multi-agent learning based on team reward.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Value-decomposition networks for cooperative multi-agent learning based on team reward

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.122003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:7fb191f3e94e5b3fe7de16cbdc57fa56778312cad074ee3d640fa2aab30b0737

Observation 6bd3e3c6-1887-42a3-b997-db2c04901e0e · outbound

This paper cites QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:12:15.515186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:5e736151d0a710cceb355f2fff08f8f4615de2eb630d28a612ecf55ab7288701

Observation c1829183-134b-4ba1-b4b0-18623b4feb9c · outbound

This paper cites Team-wise effective communication in multi- agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Team-wise effective communication in multi- agent reinforcement learning

Reference 28

Resolution
verified exact
doi, observed 2026-05-19T11:12:15.393439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:603ae402ad0dcfbfb2ca6915707076dee4003e5aaa99598d66c6bd9e0c483a51

Observation 2b4ab3f2-fa57-4167-879e-eb26ad787282 · outbound

This paper cites Scalable communication for multi- agent reinforcement learning via transformer-based email mechanism.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Scalable communication for multi- agent reinforcement learning via transformer-based email mechanism

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.124547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:43314d72f4c7248f11d19e19258ca4c90b5591122fdb5a5efb4777c92b7c36ef

Observation 8f29bd60-ff08-4799-be3c-e5f2ba6cb7c9 · outbound

This paper cites Ac2c: Adaptively controlled two-hop communication for multi-agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Ac2c: Adaptively controlled two-hop communication for multi-agent reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.129413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:f89a87f635da61a3c68d39c20653f8f8ef57f4893bded999cea50f7fd2c0efa2

Observation 681e70c1-bcf2-40cb-b08d-b11cc9476874 · outbound

This paper cites Communication in multi-agent reinforcement learning: Intention sharing.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Communication in multi-agent reinforcement learning: Intention sharing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.110986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:6aea38c072bf831f95fc6845ba673e60acc4827f332e0781552947460e227e2c

Observation f81c8048-5762-4d14-9818-05604413415f · outbound

This paper cites A sequen- tial multi-agent reinforcement learning framework for different action spaces.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A sequen- tial multi-agent reinforcement learning framework for different action spaces

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.113073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:20152df80eb8fa35163f415bbd93db286bb5902424cae0abe75004c3e779f1f7

Observation 9ec9337a-6685-4716-803a-84c6e66e4d8f · outbound

This paper cites Addressing high-dimensional continuous action space via decomposed discrete policy-critic.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Addressing high-dimensional continuous action space via decomposed discrete policy-critic

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.119805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:46e0e67301445eb169d9a99ceb2f2ebc12f312ac67e4e783d9cdae8a0b79b910

Observation d9b86bab-2a56-4e04-bbce-07778687e15f · outbound

This paper cites D- marl: A dynamic communication-based action space enhancement for multi agent reinforcement learning exploration of large scale unknown environments.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage D- marl: A dynamic communication-based action space enhancement for multi agent reinforcement learning exploration of large scale unknown environments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.127042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:6080c93a59463964f1338d694fb5de05bb5a9a02704bc377887fe05967d1476f

Observation c368e025-1df5-4fdf-bf91-ecca7d30651e · outbound

This paper cites Cooperative modular reinforcement learning for large discrete action space problem.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Cooperative modular reinforcement learning for large discrete action space problem

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.138372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:93b2754e2fd702f659cd5f58d93a039a7ebc1ec40e05a5298812dd957f64ecf5

Observation f2841c80-b544-401a-80df-7fc89409c91f · outbound

This paper cites Exploration in deep reinforcement learning: From single-agent to multiagent domain.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Exploration in deep reinforcement learning: From single-agent to multiagent domain

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.379284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:cd5f391b7c48bbaa78d8e3fefd43cd7718cc14cc211381e44f5ceb850a3365cd

Observation c5014f46-13e1-4b75-a081-c2c7d71e083e · outbound

This paper cites Rethinking exploration and experience exploitation in value-based multi-agent reinforcement learn- ing.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Rethinking exploration and experience exploitation in value-based multi-agent reinforcement learn- ing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.108434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:b2a31500dd85921bbf6556f265b8dddc7af9b14b59fb0b3685f78821ef590df7

Observation 6df1a2fa-5992-46c3-9fee-e9eb29b17a3e · outbound

This paper cites Scalable evaluation of multi-agent reinforcement learning with melting pot.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Scalable evaluation of multi-agent reinforcement learning with melting pot

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.117478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:b095e7b6e506e961137d804ff4411752123e8f35223a673b1271b77b55e54d66

Observation 34b8e147-dab5-46a9-b494-dacd214514b8 · outbound

This paper cites Sample-efficient multiagent reinforcement learning with reset replay.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Sample-efficient multiagent reinforcement learning with reset replay

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.131494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:b794263799ab9e729d96cd090b68d95fda21f6d87fdcdc7f5da8fbf93cb38485

Observation 22d7f1f8-5dbf-4f12-ad22-26194c77f6c0 · outbound

This paper cites Self- organized group for cooperative multi-agent reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Self- organized group for cooperative multi-agent reinforcement learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.168640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:dd779b66880b94b2ae006def8ff8171cd1f019729a697a168603a0828379d9c6

Observation 59a4aafd-66da-44ae-a73e-2bd64e3931bd · outbound

This paper cites Google research football: A novel reinforcement learning environment.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Google research football: A novel reinforcement learning environment

Reference 41

Resolution
verified exact
doi, observed 2026-05-19T11:12:15.375837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:fea89fa6b2c17671c84dcd4d39d51c4c90bc45f1433fe13a70bfd2f5744432e3

Observation 069fb439-390c-4f54-8954-76a4daa4e3b9 · outbound

This paper cites Pommerman: A Multi-Agent Playground.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Pommerman: A Multi-Agent Playground

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.530792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:be4101aa751eee362284f472e74671407824590ae1e4b23967516892e46e8c15

Observation b443bf14-cb00-4f18-9b96-dbd5f29398b6 · outbound

This paper cites The Neural MMO Platform for Massively Multiagent Research.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage The Neural MMO Platform for Massively Multiagent Research

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.518312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:4ec13b5d16056401f846ed48ecf75d9fa5d92b0a3b9eb8d1fe6b9e526c4e884f

Observation 0c6551ca-620b-4595-be86-ac8754e41ff0 · outbound

This paper cites Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.103922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:f99ee5c438c9c4968827ba18e2a4d7ec41d2449b5cfd93b163959e3b89e1819c

Observation 51f91da0-7c9b-4f6f-8383-20d752cbdc6e · outbound

This paper cites A survey on multi-agent reinforcement learning and its application.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A survey on multi-agent reinforcement learning and its application

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.101145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:909c6e838668a2a35b21d66de21b2abd12c413c92086af88377367659bbc0c14

Observation 8c38fe42-1ce9-4ed0-ba90-1b18238e6998 · outbound

This paper cites A Comprehensive Survey on Multi-Agent Cooperative Decision-Making: Scenarios, Approaches, Challenges and Perspectives.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A Comprehensive Survey on Multi-Agent Cooperative Decision-Making: Scenarios, Approaches, Challenges and Perspectives

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.539994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:adf9ae1fb792c0892a1508fe2a10c2f275e74cd4c82343e3956735ea0cfdd680

Observation 984784e3-839e-4323-bd40-9eabdd7c9020 · outbound

This paper cites A concise introduction to decentralized POMDPs, volume 1.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage A concise introduction to decentralized POMDPs, volume 1

Reference 47

Resolution
verified exact
doi, observed 2026-05-19T11:12:15.386684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:2f9d134854555ba01215de66ac8e409f56c4409dc7f447c33a1a98b8bb96c7a2

Observation 07c6ba2a-f762-4824-a383-bebfcc409986 · outbound

This paper cites Cooperative multi-agent control using deep reinforcement learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Cooperative multi-agent control using deep reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.115326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:db0a7b06d95cbc89e0596b599a8ca245433754f9cd689f66bed4521d4ef14f0f

Observation 936e4ac7-f1e3-440c-8576-5e3f103348af · outbound

This paper cites Multi agent deep reinforcement learning with deep q-network based energy efficiency and resource allocation in noma wireless systems.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Multi agent deep reinforcement learning with deep q-network based energy efficiency and resource allocation in noma wireless systems

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.179189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:d828c58c0b54e44b09db52964db00caf21cf4566a149f33bb6c55e3894bdae9d

Observation 9c452e70-daf3-4043-bac1-14aa5721165a · outbound

This paper cites Deep Q-Network Based Multi-agent Reinforcement Learning with Binary Action Agents.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Deep Q-Network Based Multi-agent Reinforcement Learning with Binary Action Agents

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.546177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:8867e2a35a5bc6255237c3a53726216460f4b2c0b6750352ba88a3363a3aa0ee

Observation 115426cf-df8f-4859-b2ac-899cf8841013 · outbound

This paper cites The dynamics of reinforcement learning in cooperative multiagent systems.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage The dynamics of reinforcement learning in cooperative multiagent systems

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.170507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:6156a6921388f9aa1797520be5329f70377e9c726196e17362090b016cb8f815

Observation d8cd3f2c-7f03-4ccd-a776-8a524cd3d12e · outbound

This paper cites Centralized training with decentralized exe- cution reinforcement learning for cooperative multi-agent systems with communication delay.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Centralized training with decentralized exe- cution reinforcement learning for cooperative multi-agent systems with communication delay

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.181385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:39be69cab7dd2f5462d5c08fd84bff76fe945f98b529096431575409701ddaec

Observation 1099799c-3f0b-4b0c-9b86-453c0e31561f · outbound

This paper cites An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:12:15.521735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:0e58aaa7f561a6e62d2d3b8c3dc355804ff394d7997a7cb905bfc8fb31e6b8d4

Observation d1c7f064-ba8f-4a56-af65-1fa8ab84a91a · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage The surprising effectiveness of ppo in cooperative multi-agent games

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.163669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:1b67b864e6e02665c0f06177d53bf0c19dc701cb7fd9f1009fccf3178e0dc31e

Observation 9e219d8f-4b50-4b30-95ec-2dc62869b603 · outbound

This paper cites Multi- agent actor-critic for mixed cooperative-competitive environments.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Multi- agent actor-critic for mixed cooperative-competitive environments

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.157090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:5019f82ab33ad6a3b106252ffd7fd858ac7cdf7b428526cb2776def2ebeb2aa4

Observation 91c5ef38-5182-4a1d-982f-3abd5bb0ba93 · outbound

This paper cites Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:12:15.543253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:68487533153bef66afe00b9c749719ad67e5f722e3080d6b6c1d1f4c9c46a41b

Observation ae42e91b-1b07-4bc0-b5f3-87bf3624232e · outbound

This paper cites Curriculum learning, in: Proceedings of the 26th Annual International Conference on Machine Learning, Associa- tion for Computing Machinery, New York, NY, USA.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Curriculum learning, in: Proceedings of the 26th Annual International Conference on Machine Learning, Associa- tion for Computing Machinery, New York, NY, USA

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:12:15.390481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:a185914a92eb0393e79bc32494ec95b58c4c99e0bef0cd02b50bf956f0d15190

Observation 8468a40a-ab0f-4080-921d-7b49b7d06fd8 · outbound

This paper cites Self-paced learning for latent variable models.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Self-paced learning for latent variable models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.146799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:f5d9dcd50d9fe54d241a03aaa0b51c527aa0040b5ef15538e090819ed405620f

Observation 1fe699bd-a495-4950-9b05-d576c0219a83 · outbound

This paper cites Available: https://proceedings.neurips.cc/paper files/ paper/2010/file/e57c6b956a6521b28495f2886ca0977a-Paper.pdf.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Available: https://proceedings.neurips.cc/paper files/ paper/2010/file/e57c6b956a6521b28495f2886ca0977a-Paper.pdf

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.135966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:fc6ead016f2691ec9d82012a174d3398e3408c6a8141cf13750ef57cb65e5b6b

Observation 1d11207a-4075-47ee-afdc-12c95bab9dd2 · outbound

This paper cites TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:12:15.524547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:2962550d31198ea9c18490f91c4a87ad58a887e730401953df850bc4a5a0cae1

Observation 91b17226-a61a-4b96-ac66-e242b95fd14f · outbound

This paper cites Qplex: Duplex dueling multi-agent q-learning.

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage Qplex: Duplex dueling multi-agent q-learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T11:12:16.148694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:08:51.426575Z digest=sha256:30dd9d861efba322feabde8211525374a043f6a3dc26c090323a89cc73ff1b01

Pith citing papers

No inbound Pith citation observations are available.