Pith. sign in

Paper Citation Record · LEDGER

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2502.03506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03506 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T04:16:54.807723Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact8
  • verified fuzzy20
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d4aad54-5a9c-494e-81d3-b7f10e0e85e7 · outbound

This paper cites Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:30.943737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:8f6e50a2e8130397a5816cbf48f3bad94a59e53cbb007520b94103438d06d381

Observation 28fa6271-c188-44fe-a5cc-1d07c8c45d4a · outbound

This paper cites The dynamics of reinforcement learning in cooperative multiagent systems.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning The dynamics of reinforcement learning in cooperative multiagent systems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.801505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:fb0f30d70dcb7fc9302807360ae2d48e4cf462d0b2dd701de06203a0c3ea7d64

Observation 440c6bfa-76f0-4343-82ec-77bd178c0ebd · outbound

This paper cites Stabilising expe- rience replay for deep multi-agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Stabilising expe- rience replay for deep multi-agent reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.746411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:d3135b9079afe75a5a89b9753283995771a932a0b7fe3e524d62ac25a2685d34

Observation 3e3b0115-ad2a-446b-9b21-6580b0aea4c6 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Counterfactual multi-agent policy gradients

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.785800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:4e8d76509dc49611081f7133e78da394dab35a512a556e856b36bf568de12b9d

Observation ccf0af2c-98b7-43a9-9fd1-183ba17c465e · outbound

This paper cites Cirs: Bursting filter bubbles by counterfactual interactive recommender system.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Cirs: Bursting filter bubbles by counterfactual interactive recommender system

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.789834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:a4cacedd4395d02eb491b6de6103297a330b68b16afc129910c9b956ecaa1334

Observation 503cfa3a-9b97-4e47-be70-8eec28cc4376 · outbound

This paper cites Sampling efficient deep reinforcement learning through preference- guided stochastic exploration.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Sampling efficient deep reinforcement learning through preference- guided stochastic exploration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.797050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:46f6dfde664c3bbc5027cd128008f169880b0845f585c47778ef0aa43aeee663

Observation eb549330-d5c8-4138-9a51-50d80196d2d5 · outbound

This paper cites Actor- attention-critic for multi-agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Actor- attention-critic for multi-agent reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.793459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:a6bb2c51d79f1bebe38fbf49933e169daeb7dc2dce1eb600389e37fd53af9126

Observation 2fff0a42-23cb-4797-acc6-6e8a5edbce07 · outbound

This paper cites A Maximum Mutual Information Framework for Multi-Agent Reinforcement Learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning A Maximum Mutual Information Framework for Multi-Agent Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.928004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:b2e19d6623742c81539de1d0a874d9de521b3ef03120a3fe9356e72b8882faf1

Observation 892fa9bd-5464-416a-969e-728259483c98 · outbound

This paper cites Multi-agent reinforcement learning for traffic signal control: A cooperative approach.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Multi-agent reinforcement learning for traffic signal control: A cooperative approach

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.734601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:7ffd576d91d8a59ed3a3791a675cfb08caa31ef57fb8cae139e3b0f33ad43b75

Observation aae924d3-e707-4255-91fb-fc0d536bdac8 · outbound

This paper cites A unified game- theoretic approach to multiagent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning A unified game- theoretic approach to multiagent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.781804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:185380c532445acecc8500ac7894eec9acfb47dec147d273cee0e64335b4860e

Observation 5f3e2173-aa79-455d-8cf5-0d511a204eec · outbound

This paper cites Optimistic value instructors for co- operative multi-agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Optimistic value instructors for co- operative multi-agent reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.778040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:100168e4406f9ed6501201a59a28308754048c2748794832a7cb6833266509c8

Observation baf7b414-d434-4504-abaf-5e089e76c9e6 · outbound

This paper cites Markov games as a framework for multi-agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Markov games as a framework for multi-agent reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.742587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:d4b02675d32b8f841b6e1ef41a6df36e4cab0847d1635a79483cf044b701de18

Observation 5debd818-3b4c-473a-aac7-db4c96854143 · outbound

This paper cites Cooperative exploration for multi-agent deep reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Cooperative exploration for multi-agent deep reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.774269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:bcfa10295b5eb6ff455a1e0994d3f41aefb0ebad3034fe10d279eaff351021ee

Observation aed44080-2486-40ba-ab85-1b01755c96a4 · outbound

This paper cites Multi- agent actor-critic for mixed cooperative-competitive envi- ronments.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Multi- agent actor-critic for mixed cooperative-competitive envi- ronments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.738504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:ca24978746929e02376fd1fd88df9b96e2dcc6645c3fd277727d982f617d5bf6

Observation 598393ff-2c5b-4228-b744-c825964102a1 · outbound

This paper cites Likelihood Quantile Networks for Coordinating Multi-Agent Reinforcement Learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Likelihood Quantile Networks for Coordinating Multi-Agent Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.933279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:ba585240e4380c9ab63815c6df9f5f5542d2f33561823d9f19f811804328f0ed

Observation d1922028-8854-4da9-97d6-ebce5426a694 · outbound

This paper cites Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.907248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:25da42b46e72bc0971b1b005e86b67d99c599f4a9c719973057055280bd3c2b9

Observation 7382ba68-2bbd-47a7-8697-4572e85afe5e · outbound

This paper cites Residual q-networks for value function factorizing in multiagent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Residual q-networks for value function factorizing in multiagent reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.770631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:8fb8bf7baeb1ce7fec97ac3cc690ebcb021c04637ccb9eeeb8e38a3286e9d4b6

Observation 5e9f1047-0c1b-4787-80cc-b805414b23a1 · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning The StarCraft Multi-Agent Challenge

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T04:17:30.922080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:104f2d875e6f13fa5b26fca4313d3ecadfced77d28aabea4aa246a83c35f31ae

Observation 3bb3f69d-8882-48fa-9ba7-10376295c472 · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:30.956610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:4013013d14f06d926bcdd3c26d69a232a152388971e1567ea06d1a3bf09602bc

Observation 6cf9922e-8d22-403d-a9fa-cc4706fea37d · outbound

This paper cites Resq: A residual q function-based approach for multi-agent reinforcement learning value factoriza- tion.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Resq: A residual q function-based approach for multi-agent reinforcement learning value factoriza- tion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.758213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:3f2fa315fb20edc966b7324818c6e4e01978b8cc1218f5431f8196505ff263e8

Observation 78741196-64ef-4eec-86db-bf3de7557ce7 · outbound

This paper cites Qtran: Learn- ing to factorize with transformation for cooperative multi- agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Qtran: Learn- ing to factorize with transformation for cooperative multi- agent reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.762226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:72be381af6d61ce4954eaaf43237f7ccdf4b6f9d675a1bee3031ae8c79f1fc97

Observation 58ef8e98-0c50-44e1-ab79-74985af328e2 · outbound

This paper cites Dfac framework: Factorizing the value function via quantile mixture for multi-agent distribu- tional q-learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Dfac framework: Factorizing the value function via quantile mixture for multi-agent distribu- tional q-learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.766472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:886846fcb3dc4f0ebc52defb76f6ab4b327ef4d908d9442c48def90f5f601e33

Observation 8ffdbc16-ac46-4810-a962-23a104ba1e77 · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:30.912295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:95881040aecc90845b46eb26dbe8f129593b4392985abd5ce8ce0f43ae5fbd8e

Observation 33a79463-f57b-475e-b3ad-20b001c51f8c · outbound

This paper cites Influence-Based Multi-Agent Exploration.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Influence-Based Multi-Agent Exploration

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.949668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:2f4fcd2707cab4ea7fe26649196955122cba83e513db172dd8f0b0f751ea3692

Observation 152f135b-43f8-4049-9623-59100e86958d · outbound

This paper cites QPLEX: Duplex Dueling Multi-Agent Q-Learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning QPLEX: Duplex Dueling Multi-Agent Q-Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.938525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:c3318307ed8ba6089243d5bf6902014ab3e6cd9589b76202ded392205b3ed680

Observation 47e915b9-1ee9-4749-a7a4-76832870e21d · outbound

This paper cites En- hancing collaboration in multi-agent reinforcement learn- ing with correlated trajectories.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning En- hancing collaboration in multi-agent reinforcement learn- ing with correlated trajectories

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.810003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:c766e10e019e62a9f89e23674fe84fe6940664b5e25621916ecabcd289eea763

Observation 3bf41391-8c21-4374-937f-80f6536a38ff · outbound

This paper cites Fully decentral- ized multi-agent reinforcement learning with networked agents.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Fully decentral- ized multi-agent reinforcement learning with networked agents

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.806131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:c99cb9c1d94ef334248f5d4c86d45e326ccc2c9b630a123f0c465b63327fa871

Observation 871cbf58-fcba-4061-823a-0f2f95f7812e · outbound

This paper cites Condi- tionally optimistic exploration for cooperative deep multi- agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Condi- tionally optimistic exploration for cooperative deep multi- agent reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.754518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:c3065b3717ee94570e791bb14838a9f7f241becc891d1c9bf46a05aff3317dbb

Observation a4484c9f-959e-48e1-9f12-df838671ca77 · outbound

This paper cites Qdap: Downsizing adaptive policy for cooperative multi-agent reinforcement learning.

Optimistic {\epsilon}-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning Qdap: Downsizing adaptive policy for cooperative multi-agent reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:17:31.750422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:16:54.807723Z digest=sha256:4c26b8ae82bc73a3454ddc25c32a6144f3ca20c7ba643702fbb9dc9533accdda

Pith citing papers

No inbound Pith citation observations are available.