Pith. sign in

Paper Citation Record · LEDGER

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations

As of 11 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2501.00160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00160 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:04:37.750141Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:56:58.486176Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact10
  • verified fuzzy37
  • unresolved12
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c82d0212-b734-4ce6-8b5c-1de19bd6988e · outbound

This paper cites Sutton and Andrew G.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Sutton and Andrew G

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:39.076338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.462734Z digest=sha256:19934f153791cc42ae9ceb99584195ec7e2fea46fc97b579b62648c01d44d063

Observation ede5c600-0596-4f40-916d-c7d11cdb3cba · outbound

This paper cites Q-learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Q-learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.468039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.468039Z digest=sha256:9b45ca50ca033af877157c320bb6c35e8f0724c68cbf781f8661549c000afe06

Observation fb0ee503-4963-4f04-ab85-80a77ddcb118 · outbound

This paper cites A neural substrate of prediction and reward.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A neural substrate of prediction and reward

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:39.060951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.473285Z digest=sha256:d7e622c24ea56f91395a2bf9bc0d88c2bec0d4974f399b32182b079101996d7d

Observation 2fe1db84-9e11-4071-a921-51e8eccedcf2 · outbound

This paper cites Reinforcement learning: the good, the bad and the ugly.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Reinforcement learning: the good, the bad and the ugly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:39.045745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.481195Z digest=sha256:96f68613986d048b62626fb1e38070937820d38cd536a53cee8341b285fc0848

Observation 069be5b4-5c8c-4d98-8e7a-2872ba4acde2 · outbound

This paper cites Read Montague.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Read Montague

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.486353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.486353Z digest=sha256:9bc7526982287cdae36573bb5e8696c03bd417dca3452293b9074b5721a0a05d

Observation f5710a1e-1e26-4a50-a54d-1c7fcb1720a7 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Playing Atari with Deep Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.491329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.491329Z digest=sha256:7677a14314b028e557cf6ebd59b0f496de663b181b0f30a8f90e09275c1f7b25

Observation 6d0e6ee3-847e-4b95-9d40-4f8a78af7e18 · outbound

This paper cites Human-level control through deep reinforcement learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Human-level control through deep reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:39.030325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.496707Z digest=sha256:c611734aa1da08683c31f07ddd31a17f95adee6c6651624588b3834f827e0ac3

Observation 87b19751-182e-4e01-bd46-3be7295d6753 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Mastering the game of go with deep neural networks and tree search

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:39.015489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.501092Z digest=sha256:b4068c77f8ab0961468581cb763d8be784f877d66519a613c634ef2032ee8d35

Observation 90890e8b-073a-4732-9a24-9a4103ab86fe · outbound

This paper cites Open Problems in Cooperative AI.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Open Problems in Cooperative AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.505375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.505375Z digest=sha256:6268db21071da8a21346d28a55e8778d72caf075eb5ee78f565e616f4b930e58

Observation 2f41f710-0c76-4ef8-8854-f531335c1e30 · outbound

This paper cites Cooperative ai: machines must learn to find common ground, 2021.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Cooperative ai: machines must learn to find common ground, 2021

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:39.000806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.510300Z digest=sha256:a8dda67fefb87ab81999bc923546dc5b8b91c8e7fd1760f1990b6d2b6c4abf85

Observation d0490090-1490-4591-8b81-8c9391b46e1e · outbound

This paper cites Albrecht, Filippos Christianos, and Lukas Sch¨ afer.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Albrecht, Filippos Christianos, and Lukas Sch¨ afer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.985918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.514752Z digest=sha256:89d345c740e372eb27688b3dc1665722a57a0058fc8bf377d5828979cd5c14ec

Observation 66e493d2-750a-4f41-8a3f-b583ca55f98b · outbound

This paper cites Multi-agent reinforcement learning: independent vs.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Multi-agent reinforcement learning: independent vs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.971316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.519210Z digest=sha256:8884119726abc1844fbb540c0f4ef92f4b5fb7b356a09885bf2d039cadce835a

Observation 074fdb90-989e-4979-9108-594560fbbae3 · outbound

This paper cites A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.523950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.523950Z digest=sha256:4d1089a31a112aff5467b37db79e670c0ad7ec455b9056ef4b27e181f13d6b34

Observation 5927b641-37e8-4527-a776-0eb37471a74f · outbound

This paper cites Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.956202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.528783Z digest=sha256:446fa0866dd47843c1635545087dce5f6a2d4a7c6369faf2016930b11baf436b

Observation 888d0c23-4265-4fc3-9858-4795c1980cd8 · outbound

This paper cites A survey and critique of multiagent deep reinforcement learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A survey and critique of multiagent deep reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.940143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.533148Z digest=sha256:e891367ee4db075b2f06b0a725d2cb61e32643e51a5e53e92bd8c066f0a8740a

Observation 431129d0-3e96-4b6d-aeea-44cd422c9dfd · outbound

This paper cites Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.538087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.538087Z digest=sha256:1dd45687f1738fea440936198dd4e0e151917dd945b0ec4326ddfba315fc3891

Observation 0687e0cd-9be0-450a-bcb2-dda5084f9e2d · outbound

This paper cites Learning through reinforcement and replicator dynam- ics.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Learning through reinforcement and replicator dynam- ics

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:04:38.350491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.543004Z digest=sha256:858314d31fbe9ec6427b8d19214ad15f9d3ed3c3c1377048e442e228610fab54

Observation 755cd89e-5898-46fc-b5c3-189ffc5fdf46 · outbound

This paper cites A selection-mutation model for q-learning in multi-agent systems.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A selection-mutation model for q-learning in multi-agent systems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.925545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.547413Z digest=sha256:27cc4ab8a05d298457e93bd168eef3a686fa5cc911d5460ed5ae8287c0ba1dcf

Observation e0a297b2-13ba-4172-a790-f3315cba42ae · outbound

This paper cites Doyne Farmer.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Doyne Farmer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.556134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.556134Z digest=sha256:66cb2f202dcd409ce61b4f99227ba4bcb1a3b867ca03223094205431de021088

Observation 05adf438-9b55-4cb9-8e4a-8826f82021a7 · outbound

This paper cites Crutchfield.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Crutchfield

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.560867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.560867Z digest=sha256:cd0477975381ff37c4164264ec5f2e34fcf64c41ebbe86bdcadb2b220a3a47ff

Observation 608361af-33bd-4ec0-ab50-0325dfeb9be6 · outbound

This paper cites Crutchfield.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Crutchfield

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.910090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.565691Z digest=sha256:36d918babfc95003e02523fe8e1ac278dcadf3b79502738387e7e6f6d391bc94

Observation 818a8dbe-0bb5-4302-9f5a-d6726841646b · outbound

This paper cites Individual q-learning in normal form games.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Individual q-learning in normal form games

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.894600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.574579Z digest=sha256:539e9cd20458b44732618589e5aa783c0e50fda4ae043dc1ee6a0e665a02598c

Observation 7c5e8bd7-72b2-4ec2-b3fb-f4fd290f5608 · outbound

This paper cites Reinforcement learning dynamics in social dilemmas.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Reinforcement learning dynamics in social dilemmas

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.881520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.579301Z digest=sha256:9acc14ff085471a306b9f40374befcebaf465b81af8bce2812e262974d2091a9

Observation eb3d957f-51e8-4d4a-88f3-1209bb6600f9 · outbound

This paper cites Learning and equilibrium.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Learning and equilibrium

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.867165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.583929Z digest=sha256:d56bb75f1194ad9d0c560436a685c361875e4203351f343099413f1d54f940ab

Observation 53a100fc-dd1d-4a89-8e68-6773734a60c6 · outbound

This paper cites Intrinsic noise in game dynamical learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Intrinsic noise in game dynamical learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.851549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.588353Z digest=sha256:d9bc6a53c5199013532b2e806e9d35ccc2eb6f8ee8c37a26d1c80b0ae2e6ded9

Observation 66a06a5b-048b-4f66-a589-3eb444863253 · outbound

This paper cites A theoretical analysis of temporal difference learning in the iterated prisoner’s dilemma game.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A theoretical analysis of temporal difference learning in the iterated prisoner’s dilemma game

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.835719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.598200Z digest=sha256:fa23546c432605e2a7e46a77a771ca2c8d1121c021eb5568e8ffc13e9304a3d1

Observation d596418a-fa44-4095-b7e4-b2ac82b9d517 · outbound

This paper cites Classes of multiagent q-learning dynamics with epsilon-greedy exploration.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Classes of multiagent q-learning dynamics with epsilon-greedy exploration

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.820001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.602787Z digest=sha256:1d05174332929dff90bda94844ce4d715fa950028d8239ea129aa27b655ae366

Observation a1b6a395-0b89-4a1e-bb55-6b32f1bdfa33 · outbound

This paper cites Numerical analysis of a reinforcement learning model with the dynamic aspiration level in the iterated prisoner’s dilemma.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Numerical analysis of a reinforcement learning model with the dynamic aspiration level in the iterated prisoner’s dilemma

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.805147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.607193Z digest=sha256:70e63d9a4d8c5c0866da984a7c748383dd8dc373ffa84d38d9e2b0b8a2da3ad6

Observation 7dbba3b5-f465-4bde-8d67-01301685e1b5 · outbound

This paper cites Cycles of cooperation and defection in imperfect learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Cycles of cooperation and defection in imperfect learning

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-10T23:04:37.611648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.611648Z digest=sha256:e0201b9136e8dac7b7bdb9904ef493653cae71421f784683148d7aa606915083

Observation f1953e47-2139-4c63-bb1c-109378b2a46e · outbound

This paper cites Dynamics of boltzmann q learning in two-player two-action games.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Dynamics of boltzmann q learning in two-player two-action games

Reference 30

Resolution
verified exact
doi, observed 2026-08-10T23:04:37.865856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.616159Z digest=sha256:2b5afd12f3a065ba959167bae42797555e8efd4d78f3edeaca94a9d00e9b3e3e

Observation 6c92b88c-3c2f-4dde-9e7b-347a99363090 · outbound

This paper cites Continuous strategy replicator dynamics for multi-agent q-learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Continuous strategy replicator dynamics for multi-agent q-learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.790443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.620816Z digest=sha256:9678ec4e0eab942965c096e6683645923c1f0a937aea691a07494324eef39d22

Observation c6866ed2-8ffd-48fc-b9ab-ea5b82e58334 · outbound

This paper cites Doyne Farmer.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Doyne Farmer

Reference 32

Resolution
malformed identifier
no resolver link, observed 2026-08-10T23:04:37.625483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.625483Z digest=sha256:7f0bbdb2177c9a332d01b5571f33f8d7df4468b56bd5309494dedd94a64e3ae6

Observation 32690eba-1bac-4a08-8285-be07bfd8d0ea · outbound

This paper cites Evolutionary dynamics of multi-agent learning: A survey.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Evolutionary dynamics of multi-agent learning: A survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.776330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.630631Z digest=sha256:7f34e175c860e30598535623199afaaaaec2ee33d2db5eb86ed8e103903b7a68

Observation c8f3080a-3b9c-44d5-8592-53aca54a9702 · outbound

This paper cites an unresolved cited work.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.635376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.635376Z digest=sha256:f853c9bfbae4f97857d03e2577f61b27830d61e13bd7dc59f6c669179a0d8aac

Observation 0b931376-2dbd-4507-9242-c7bdc6a05b25 · outbound

This paper cites Donges, and J¨ urgen Kurths.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Donges, and J¨ urgen Kurths

Reference 35

Resolution
verified exact
doi, observed 2026-08-10T23:04:37.830920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.640348Z digest=sha256:fbe24894d88618d09a341b4177cc1e3057bcee0432ffdef29c0e97cbcf3bf89a

Observation bc80f7e8-be7f-49e0-80b8-70a8c21ad044 · outbound

This paper cites Modelling the dynamics of multiagent q-learning in repeated symmetric games: a mean field theoretic approach.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Modelling the dynamics of multiagent q-learning in repeated symmetric games: a mean field theoretic approach

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.762892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.644820Z digest=sha256:e715cd832a78180e468f5caeb88b12b616e6bf3fbc5926c3e1c45bd291093255

Observation d170359c-ecde-4b7a-b828-3f5ba27c1ee3 · outbound

This paper cites Dynamical systems as a level of cognitive analysis of multi-agent learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Dynamical systems as a level of cognitive analysis of multi-agent learning

Reference 37

Resolution
verified exact
doi, observed 2026-08-10T23:04:37.816019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.649221Z digest=sha256:f9f7a830d20bd08494f5fea7c9edf58f1444c34b7298a5bea6f39a37b03d6d9b

Observation d5c323d6-1f11-4456-b750-81e19428dd89 · outbound

This paper cites The Dynamics of Q-learning in Population Games: a Physics-Inspired Continuity Equation Model.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations The Dynamics of Q-learning in Population Games: a Physics-Inspired Continuity Equation Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.653856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.653856Z digest=sha256:efcf32372c228fee598e07e20fecf9d0bb132f703ce04fcbfae033d0f7993e24

Observation 464b5458-f6ed-43d9-a852-0e76a6dc80f2 · outbound

This paper cites A formal model for multiagent q-learning dynamics on regular graphs.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A formal model for multiagent q-learning dynamics on regular graphs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.748897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.658248Z digest=sha256:16217637a79a90cafde484a5a642bee8dd0508c48fb0c725ca5364817c9f1b13

Observation 83891b36-205b-4e5b-8b72-46159c69cfd1 · outbound

This paper cites Modeling the effects of environmental and perceptual uncertainty using deterministic reinforcement learning dynamics with partial observability.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Modeling the effects of environmental and perceptual uncertainty using deterministic reinforcement learning dynamics with partial observability

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.734983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.662675Z digest=sha256:68cbc23d43814c0ec1aaea7308cca9b37043f552db44f5cdfb666e933b9fc96a

Observation ae5d1519-d53c-4637-a811-f02ce349dffd · outbound

This paper cites Exploration-exploitation in multi-agent learning: Catastrophe theory meets game theory.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Exploration-exploitation in multi-agent learning: Catastrophe theory meets game theory

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.718840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.666984Z digest=sha256:ead44536978be88c450453ffd2dfddcf4a12f84496ec3ee5e8f0905b6fc92781

Observation b204136b-54d9-4b1b-bf93-c084026b2a45 · outbound

This paper cites Frequency adjusted multi-agent q-learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Frequency adjusted multi-agent q-learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.703830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.670922Z digest=sha256:f85dfd1eb549c43b37ba0f0ed7cb3529dbd9a73a607daec5fac4c6c1c56d7cb9

Observation dde5802c-95a4-4165-9a28-757e77cfda30 · outbound

This paper cites Evolutionary Multi-agent Reinforcement Learning in Group Social Dilemmas.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Evolutionary Multi-agent Reinforcement Learning in Group Social Dilemmas

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:04:38.157481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.675296Z digest=sha256:7696c1b68b2cd8f965b7f4d2953ccb750bd447309187d6fc5c79053079a5eae5

Observation 007617c1-8e8b-4f8c-ab5a-06becd7747ec · outbound

This paper cites Sandholm and Robert H.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Sandholm and Robert H

Reference 44

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:04:38.136190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.680420Z digest=sha256:82515743b4e7f70b451988517dd8c3ebeefcaae1ed13ef5b99105e81342bbb22

Observation ba7a4dca-6d42-4315-b8ab-b6ca5c79b24e · outbound

This paper cites Faq-learning in matrix games: Demonstrating convergence near nash equilibria, and bifurcation of attractors in the battle of sexes.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Faq-learning in matrix games: Demonstrating convergence near nash equilibria, and bifurcation of attractors in the battle of sexes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.689075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.684780Z digest=sha256:9ea8a152ac0ef9036700504e4440d5c13884fc0b686f2926c7d0991d053bc607

Observation cf2886f0-3d29-4440-84c1-78048f471843 · outbound

This paper cites L´ evy noise promotes cooperation in the prisoner’s dilemma game with reinforcement learning.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations L´ evy noise promotes cooperation in the prisoner’s dilemma game with reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.673951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.689137Z digest=sha256:cf22b6c775803097edb4936283ab22e87eab725e0abc93f9472162a26b97f560

Observation da0760aa-9940-4be1-b097-a0450dcfaa2c · outbound

This paper cites Limiting dynamics for q-learning with memory one in symmetric two-player, two-action games.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Limiting dynamics for q-learning with memory one in symmetric two-player, two-action games

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.658704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.693543Z digest=sha256:0c3c667c28c090e126312eec471a920067d1b6802a848950236fef3df60213bc

Observation 2c497bf0-d8f4-45a3-b1dd-2ed07bd7b9fc · outbound

This paper cites Q-learners can provably collude in the iterated prisoner’s dilemma.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Q-learners can provably collude in the iterated prisoner’s dilemma

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:04:38.040057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.698125Z digest=sha256:7943eebf592c04f3b77adb6aa2d0f36adc9ffc66990d8b5471506c56bf420d9f

Observation 87679894-bb4d-4df0-be0a-d0ff9041c445 · outbound

This paper cites Symmetric equilibrium of multi-agent reinforcement learning in repeated prisoner’s dilemma.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Symmetric equilibrium of multi-agent reinforcement learning in repeated prisoner’s dilemma

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.643861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.702473Z digest=sha256:cfe724a3d998794337acb456561b2e623bf14b0372f1a5f33c4a53e9a93b27c1

Observation 7f307846-458d-4acd-9eab-b0e27043a061 · outbound

This paper cites Q-learning in two-player two-action games.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Q-learning in two-player two-action games

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.627730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.707053Z digest=sha256:193ec3ca9122f311c55bb2593ae2cceaa63db82cc9c6bfbe57307dc2bccbe2b9

Observation af5294dc-6586-4fa4-bb30-47d8b12cac0a · outbound

This paper cites Melioration learning in iterated public goods games: The impact of exploratory noise.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Melioration learning in iterated public goods games: The impact of exploratory noise

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.612357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.711573Z digest=sha256:ec0e4ab31684718e2f94665a28354464d90b0ac3771d4ecab3784ebb9344f734

Observation 6dcea1a7-2004-4da3-8d4d-1e2f3c91b413 · outbound

This paper cites Rein- forcement learning and decision making in monkeys during a competitive game.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Rein- forcement learning and decision making in monkeys during a competitive game

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.595972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.715929Z digest=sha256:5cd7dce65076c7c1b91e11ac092db51c96be61c9d983c998e6a6c5a229b63bf1

Observation b524ab3f-a1ab-4c81-a39a-8e307a7893e1 · outbound

This paper cites Valuation of uncertain and delayed rewards in primate prefrontal cortex.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Valuation of uncertain and delayed rewards in primate prefrontal cortex

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.578784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.720573Z digest=sha256:39852ff1bcdfa2031b99911b7e8c7635c1af00c2dd56afa7a1a5aa040cc080cb

Observation b408edeb-f28f-4075-a9c9-697d34c90c7e · outbound

This paper cites an unresolved cited work.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-10T23:04:37.798611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.725633Z digest=sha256:9a73bab52ee225bc188f955337c2c8e2acdb8e31a5f6fe87dcb1d8de85b39514

Observation 4420efb1-5bf7-463b-983d-58b89ed118c1 · outbound

This paper cites Batch Reinforcement Learning, pages 45–73.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Batch Reinforcement Learning, pages 45–73

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:37.730146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:37.730146Z digest=sha256:e96815df5847a0d3c39db4227cf996f7b5c6f57e27a0a17e5608aabb011f3273

Observation d390adb9-6282-47e6-955c-0976dca97129 · outbound

This paper cites Quantal response equilibria for normal form games.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Quantal response equilibria for normal form games

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.564275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.735267Z digest=sha256:00e50ec35e524b7357883dbf8d1a364c8751b34bbd37ca5093f50cf17a0f33cc

Observation b01182b5-d586-484e-8dfb-5fefa770ff00 · outbound

This paper cites High-stakes failures of backward induction.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations High-stakes failures of backward induction

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.549040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.740001Z digest=sha256:8252bee3caf84c55e2606b1d2aa8c8730ecdb78818bfb24dd80d89f42697fac3

Observation 00597c54-f9fb-4af8-8538-a57afa102217 · outbound

This paper cites Timing of transients: quantifying reaching times and transient behavior in complex systems.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Timing of transients: quantifying reaching times and transient behavior in complex systems

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:04:38.533809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.744647Z digest=sha256:59089edadbc58ae436ef08963f35b8222ead14651b2e214433cee2714ca65516

Observation 334dd36e-5558-4080-b6d5-f283cc63ff50 · outbound

This paper cites an unresolved cited work.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:04:38.517277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.750141Z digest=sha256:6776331db95d4b9bdaf04041a5bf9d32eff11327e04dbb187a87bd9270cb761a

Observation f480cf6d-f0b7-465a-b8ee-bc4823072742 · outbound

This paper cites ISBN 1581136838.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations ISBN 1581136838

Reference 2003

Resolution
metadata mismatch
raw_fallback, observed 2026-08-10T23:04:38.269857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.552134Z digest=sha256:2b18b7d38b4febf8f4e811cda62e13c3a28b3cb8510081a2a5c8ff588b23ff46

Observation 302cb385-d681-445d-a95a-a63c593f0215 · outbound

This paper cites URL https://link.aps.org/doi/10.1103/ PhysRevLett.103.198702.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations URL https://link.aps.org/doi/10.1103/ PhysRevLett.103.198702

Reference 2009

Resolution
verified exact
doi, observed 2026-08-10T23:04:37.889657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.592510Z digest=sha256:c42c2ad2b7735b73f28c2c50a623abcdd32cb62ba7c69fb0015d38bfa39697a4

Observation 56916ade-5a51-445c-a2c6-1b176da410fa · outbound

This paper cites URL http://dx.doi.org/10.1016/j.physd.2005.

Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations URL http://dx.doi.org/10.1016/j.physd.2005

Reference 2789

Resolution
verified exact
doi, observed 2026-08-10T23:04:37.903873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:04:37.570146Z digest=sha256:09ae4908d595032a8a400b7b1e136f606ad611577dd6542017ae976efd689788

Pith citing papers

Observation f8fcddfd-3b42-4b8a-b4e6-acceba4a25ad · inbound

An Agent-Centric Dynamical Systems Perspective on Multi-Agent Reinforcement Learning cites this paper.

An Agent-Centric Dynamical Systems Perspective on Multi-Agent Reinforcement Learning Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:56:58.486176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:56:58.486176Z digest=sha256:b9694f4f122dc2a3f39eddf1194a7bc779dae5ad8a61099487b261d3bd2903fb