Pith. sign in

Paper Citation Record · LEDGER

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2411.11099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11099 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:02:50.937962Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact5
  • verified fuzzy24
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72d251e7-2bb1-4349-9f65-efd38b2e125a · outbound

This paper cites Decentralized multi-agent deep reinforcement learning in swarms of drones for flood monitoring.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Decentralized multi-agent deep reinforcement learning in swarms of drones for flood monitoring

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.452627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.816034Z digest=sha256:3cfeea8599f8e76a042fa482a592c2af52e620a56b4de4e5773ffc24028cf289

Observation 57261a35-1a77-42b1-a1f8-afc421fc9a77 · outbound

This paper cites Decentralized control of quadrotor swarms with end-to-end deep reinforcement learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Decentralized control of quadrotor swarms with end-to-end deep reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.445297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.819467Z digest=sha256:4b7382e06fb30eeee5f1c0bd8149c0a223f73ec3ad29539faf4dec78b0e3e7bb

Observation 3d477659-2268-42e6-9aa9-59a678f415a2 · outbound

This paper cites Opportunities for multiagent systems and multiagent reinforcement learning in traffic control.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Opportunities for multiagent systems and multiagent reinforcement learning in traffic control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.437131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.822233Z digest=sha256:7cc01ddba8fcdccd7bf1552380eb6b59d780a669004809ecd32310ceba89e36a

Observation f1fe5dcb-5ac1-47db-8a02-1a0b1c86d927 · outbound

This paper cites Superhuman ai for heads-up no-limit poker: Libratus beats top professionals.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Superhuman ai for heads-up no-limit poker: Libratus beats top professionals

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.825257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.825257Z digest=sha256:975fce3829217f71f4c3c1a729908e744b6c56bcf5b6e7434110d848854f0445

Observation 26b1deee-693b-4e3b-8f42-3461600259e3 · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.828554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.828554Z digest=sha256:e618b69f589f8bdfa903b2024698b7857cef6251a09bf8ae8fc09e272cb159bb

Observation d05c8313-ede0-4df6-bfe9-d8ed70b62c87 · outbound

This paper cites FACMAC: Factored Multi-Agent Centralised Policy Gradients.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning FACMAC: Factored Multi-Agent Centralised Policy Gradients

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.831537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.831537Z digest=sha256:c12aa26d6776c15179c2fc280ca91d08c89106c94e9af0f4cc5c0505df67ccb4

Observation 2bc4d3b8-fecd-41aa-9329-0572a11feec0 · outbound

This paper cites Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.834956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.834956Z digest=sha256:f1aece7e265a9b1095b5aea885fff9dbfefdc1c1bda28a87cef9632358005206

Observation 9f2a5d52-d1ab-4360-997b-3474b1b5178b · outbound

This paper cites Learning to communicate with deep multi-agent reinforcement learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Learning to communicate with deep multi-agent reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.838105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.838105Z digest=sha256:1e087b5831186d6bd5bb9b97e9e8f0c05775f0e06117a43b2d83b771f4c00c12

Observation 253004a3-3f35-48b4-8924-eea47b4436e7 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Addressing function approximation error in actor-critic methods

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.840726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.840726Z digest=sha256:640c50fa0c10a2bb942432a6bd3420b7af9d260f449bf6405e16dc7e29b42d40

Observation b09c1bca-cda0-4783-bced-1b1e04acfc49 · outbound

This paper cites Uneven: Universal value exploration for multi-agent reinforcement learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Uneven: Universal value exploration for multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.414338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.843311Z digest=sha256:4f803b4db99c4fd6beb77004d03f695450cbb76bca0bf9cb140cb3b8873c2c9f

Observation 66319e40-e65c-4364-aab4-73d23abc5cb6 · outbound

This paper cites A survey of learning in multiagent environments: Dealing with non-stationarity.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning A survey of learning in multiagent environments: Dealing with non-stationarity

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.405993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.845499Z digest=sha256:dfae60a179add97a18e9d158dff518b36317261c7fa21328081e96c18f8561cd

Observation d7e7c90c-d383-4bd5-9e30-fee9876f8e8d · outbound

This paper cites I2q: A fully decentralized q-learning algorithm.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning I2q: A fully decentralized q-learning algorithm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.397683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.847732Z digest=sha256:d8431e4cacc1ce5f0ca9ede2cea3c5f9eb3c707b5590d70be73f07c64b46ebac

Observation 5c411b02-7e39-4cb0-bf6b-d0a2aada98d7 · outbound

This paper cites What uncertainties do we need in bayesian deep learning for computer vision? Advances in neural information processing systems, 30, 2017.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning What uncertainties do we need in bayesian deep learning for computer vision? Advances in neural information processing systems, 30, 2017

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.849913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.849913Z digest=sha256:a89db68ea5c961903d21bfd3f20220ae72ae1f3b3a15d934ad159d98acca4727

Observation 12cca604-5929-4333-9beb-2c192a5ffc4d · outbound

This paper cites Simple and scalable predictive uncertainty estimation using deep ensembles.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Simple and scalable predictive uncertainty estimation using deep ensembles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.852126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.852126Z digest=sha256:02d3d6ed0e91bee5fcdf8cd24e45de57400adce50aa87452ecaf12abef1a1b8f

Observation d6d90c6b-4a4d-4bea-9318-a30f84f3b249 · outbound

This paper cites An algorithm for distributed reinforcement learning in cooperative multi-agent systems.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning An algorithm for distributed reinforcement learning in cooperative multi-agent systems

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.380153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.854498Z digest=sha256:be552bd850ba35604161813b970d604b3eaac6adab27f9ed347b9edcaf2aa31f

Observation 21c22cbf-69d6-47ac-b469-763ee94c2cd0 · outbound

This paper cites Solving homogeneous and heterogeneous cooperative tasks with greedy sequential execution.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Solving homogeneous and heterogeneous cooperative tasks with greedy sequential execution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.370306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.857146Z digest=sha256:d0726aac291b5ca8445d1d3586da940b915de3d5acb27fb2825b26c00e419584

Observation a850375b-5929-4165-a91b-7639c631c786 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.859294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.859294Z digest=sha256:316307c3296cb9ffc257305cf25daea220da4b630a180280e96152d6785a8702

Observation db37c27d-fe6f-4d6b-a197-84f7a2344afc · outbound

This paper cites Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.356655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.861937Z digest=sha256:d1cf86bf6de3f0def6e07f0489b8a67cb703ad7c0a2cf91bd9bc9c6a6cd7b417

Observation 7c248329-1e89-43f5-89f0-38ba86288171 · outbound

This paper cites Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.348192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.864580Z digest=sha256:10df6c76e768be2764786f9470bca90fccfba2638b7848d14e1deaefbf67d6eb

Observation 701cad7e-da14-4fe5-ac1e-deea4d84aed4 · outbound

This paper cites Deep exploration via bootstrapped dqn.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Deep exploration via bootstrapped dqn

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.339002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.869836Z digest=sha256:71beb3794bcf254ba7a2372de7ecac80f8b3dce08336e3412292623a1ecd89cc

Observation b301c3ff-e77a-4cbe-bccb-d715a47f91ac · outbound

This paper cites Negative Update Intervals in Deep Multi-Agent Reinforcement Learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Negative Update Intervals in Deep Multi-Agent Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:02:51.181171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.872930Z digest=sha256:d9e4cfcb5b1191df443995af965f1f5cceb29ea481e781af3638c1a2a6d2e010

Observation e39c91a2-a02c-4551-a808-27884d86b190 · outbound

This paper cites The analysis and design of concurrent learning algorithms for cooperative multiagent systems.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning The analysis and design of concurrent learning algorithms for cooperative multiagent systems

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.330797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.876083Z digest=sha256:4f0c3f0e9593c6138099a6580e3c2caf69d7dfb2fbf20f1c1cd42fd14cc5a9b3

Observation c73bb3eb-2cf8-4b5e-8a27-9cca2c31ce0e · outbound

This paper cites Biasing coevolutionary search for optimal multiagent behaviors.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Biasing coevolutionary search for optimal multiagent behaviors

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.322597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.878684Z digest=sha256:098b2a5742e50c6367f2d5e71dba14713787d62b76f10e23ce273930457a4efb

Observation 9dbfb156-17b4-4692-9c2b-732954e7d223 · outbound

This paper cites QMIX : Monotonic value function factorisation for deep multi-agent reinforcement learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning QMIX : Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.314344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.881343Z digest=sha256:15b1d62b687b6df1e9517bb3963254e5f71b9267ba49fd6620114bc56a3dde3c

Observation 94d80098-1ecc-42c1-8edd-4dbcab3f8bc3 · outbound

This paper cites Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.305956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.883933Z digest=sha256:b87c051c7a2aa517d581be9fde1a81484aeaf4d9b6ea644109c4e92aae0de0fc

Observation 5c9c8afc-ee6d-4c4f-8a05-f71ba4094d66 · outbound

This paper cites Monte Carlo statistical methods, volume 2.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Monte Carlo statistical methods, volume 2

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.886689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.886689Z digest=sha256:7758ec26d3a8215271c7212aaf9dfbf54f7d806cdb50e2321c1d2d305d1bf020

Observation d9f1adbf-ade2-4182-9951-8e1ebbc70782 · outbound

This paper cites Sahraoui, M.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Sahraoui, M

Reference 27

Resolution
verified exact
doi, observed 2026-08-12T19:02:50.960269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.889334Z digest=sha256:dd28d1c0cdba6146f9486e0976425e259eac8c950af9e4f122df8d0dc0dc1053

Observation c766bd39-b8c6-4f60-9cc6-262f82d7d60e · outbound

This paper cites Planning to explore via self-supervised world models.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Planning to explore via self-supervised world models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.292694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.892203Z digest=sha256:357ac10d3fbb3c5af7ab07056566fdb3dbcf872a9f36f2f0c2c40020345f7892

Observation b658dd84-9677-499e-bdec-cc87e4438fda · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.894967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.894967Z digest=sha256:dccf4f87fd0353585bceeea82d8e5dc40d66dbaaa0d758cedf1dd6990ba51203

Observation e89b844a-6ce9-48c8-89dc-13ba3e2f842b · outbound

This paper cites CURO: Curriculum Learning for Relative Overgeneralization.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning CURO: Curriculum Learning for Relative Overgeneralization

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:02:51.162373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.898138Z digest=sha256:80a4c7b07accbe4dfbc4eaba68c5f31923f10d907ac5d79b6dea624c3c529b61

Observation e58a1369-5a43-4c75-8588-7ed33cf1fc03 · outbound

This paper cites Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.900982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.900982Z digest=sha256:300fc50c769a44a410a6049ff7382de0a5a9703cc0cf44c03b0afa24a1bf17cb

Observation 6257f265-875c-483e-aca8-2c1362057803 · outbound

This paper cites Lectures on parametric optimization: An introduction.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Lectures on parametric optimization: An introduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.903553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.903553Z digest=sha256:0481fbd8d831589ecc45faf33382e16cff235cd67f8b1db73d1c8a4f794362ca

Observation f49fef85-120a-4fa9-8d88-4b18567286a0 · outbound

This paper cites A fully decentralized surrogate for multi-agent policy optimization.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning A fully decentralized surrogate for multi-agent policy optimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.273828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.906223Z digest=sha256:ceca40bb2e4206e3413634b7e6cbce613f015cff66a564719e5cb12fccd6629f

Observation fa267891-e6d2-41ea-b754-9f44bec6948d · outbound

This paper cites Exploit reward shifting in value-based deep-rl: Optimistic curiosity-based exploration and conservative exploitation via linear reward shaping.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Exploit reward shifting in value-based deep-rl: Optimistic curiosity-based exploration and conservative exploitation via linear reward shaping

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.263879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.908920Z digest=sha256:2ed0c7467d64805f7d9d266940e29243c39b0465768a9cbebb75d6220766e9ae

Observation 18c1f776-6e84-4c0c-aa5e-71197a4dd426 · outbound

This paper cites Multi-agent reinforcement learning: Independent vs.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Multi-agent reinforcement learning: Independent vs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.253083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.912537Z digest=sha256:9fc8ad7dbc9f9ba9dbd3dc86ba2d0f40449de721a26f68e39d167a25bf8329f2

Observation 23201714-95c8-48f9-b119-170794099c0f · outbound

This paper cites Multi-agent deep reinforcement learning-based trajectory planning for multi-uav assisted mobile edge computing.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Multi-agent deep reinforcement learning-based trajectory planning for multi-uav assisted mobile edge computing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.244685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.915438Z digest=sha256:bb63cc912c15a6b000553f61cb7411be67bb88eb41b77cc79bfd2b6ad611b4f8

Observation 9debfa1b-c21c-4a10-ab2f-55d3760cef5a · outbound

This paper cites Lenient learning in independent-learner stochastic cooperative games.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Lenient learning in independent-learner stochastic cooperative games

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.236459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.918218Z digest=sha256:70f06075b82b939a5c994e5e92bc33fd534f73cd51a38c3f28a6aaf6d52208c5

Observation 2a0cc246-8bda-41a2-9881-b2d6d38cdd58 · outbound

This paper cites Multiagent Soft Q-Learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Multiagent Soft Q-Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:02:51.150058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.920797Z digest=sha256:7c13379804d93ac5ca1e71a6082578775c58b1cbf06fe92eb819d3e2f45fb02a

Observation 0e30e6ec-ad04-41e5-be26-db7c23e96791 · outbound

This paper cites An analysis of cooperative coevolutionary algorithms.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning An analysis of cooperative coevolutionary algorithms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.228257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.923914Z digest=sha256:2d4d00feba544e291d2b8bbab0f7956acfb49d1482e6d97bec6ac2451b31a142

Observation c7714f3f-054e-41d6-b19f-15301083ebf4 · outbound

This paper cites Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.926448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.926448Z digest=sha256:309a97714c7efa3b86b4eadd59cd0ada986e49a6d082d6fda5ce9a560266dc6c

Observation baad5374-9635-4b6f-80d0-27b18ba0bbbc · outbound

This paper cites Learning multi-agent coordination for enhancing target coverage in directional sensor networks.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Learning multi-agent coordination for enhancing target coverage in directional sensor networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.220966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.929065Z digest=sha256:2f3001dcb378c06de54ebdf79382d33c4394b69119f628ea496f36d007f95677

Observation 6d5c0fa9-3be5-4485-8e7b-c5b38728aa6b · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning The surprising effectiveness of ppo in cooperative multi-agent games

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.931292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.931292Z digest=sha256:e5bab64acebfcda2a56e53dbb50f513cf88495aa27ade4919d5df8a115b166ea

Observation 6257f696-f919-422d-a40c-a81bc2672f00 · outbound

This paper cites an unresolved cited work.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Unresolved cited work

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-12T19:02:51.131717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.933532Z digest=sha256:0538a3c7c4b149b43822aa67de4036b569f1bf0a404999f2b54968006120b6cd

Observation 161b8611-60a8-4713-b66f-16235487f7b7 · outbound

This paper cites Smarts: An open-source scalable multi-agent rl training school for autonomous driving.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning Smarts: An open-source scalable multi-agent rl training school for autonomous driving

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:02:51.207029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T19:02:50.935751Z digest=sha256:4e4b1b552a13fbdc60062e33157cb0a8255e742c1f41aaf8ad7cea6a33b756b6

Observation 15ce17ea-bca8-44f1-92d6-9bbf30194d75 · outbound

This paper cites A Survey of Multi-Agent Deep Reinforcement Learning with Communication.

Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning A Survey of Multi-Agent Deep Reinforcement Learning with Communication

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:50.937962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:50.937962Z digest=sha256:faae86686e491e24a7a1cd384240b4f159c2c625acd1007864bdd86a8564c650

Pith citing papers

No inbound Pith citation observations are available.