Pith. sign in

Paper Citation Record · LEDGER

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework

As of 18 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.15856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15856 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:34:32.435835Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f531a07-aa77-45b2-b60b-98872ecd42d3 · outbound

This paper cites Threshold bandits, with and without cen- sored feedback.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Threshold bandits, with and without cen- sored feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.962137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.291625Z digest=sha256:af63d1e2a05692d4d7eb307601755569c7518dea794fa838ea90ce765f28a46f

Observation e3315ede-c7fb-4eaa-813a-db83957c31df · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Asymptotically efficient adaptive allocation rules

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.759844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.370030Z digest=sha256:0a4d4f8a4ac684c8960372221eb881387da85d3d260b97a0dd8d2e6411caed0c

Observation 8665debe-a4a6-4909-9f16-b3ce5c53a68b · outbound

This paper cites On distributed cooperative decision-making in multiarmed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework On distributed cooperative decision-making in multiarmed bandits

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.717408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.381352Z digest=sha256:e002852d4e6a398602580f61724234946675edb63b002f6fc59c644c48be51f9

Observation b747bd04-cc0f-4a93-b293-898c383c476c · outbound

This paper cites Social imitation in coopera- tive multiarmed bandits: Partition-based algorithms with strictly local information.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Social imitation in coopera- tive multiarmed bandits: Partition-based algorithms with strictly local information

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.699203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.386557Z digest=sha256:e97988e3ad9e2da74271654bee87a4d62c821ad731a6ce3bd149b93ec195afeb

Observation cd592c1f-edc7-412c-a330-3b2c5d319825 · outbound

This paper cites Bandit algorithms.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Bandit algorithms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.662968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.397080Z digest=sha256:372fe26fd4a7bd6ecd0fa0e1aaedefee3b170895b11f68a7b04b14422535b6e8

Observation 5c40fdaa-0b59-461e-84b6-268bd0b0a603 · outbound

This paper cites Decentralized Cooperative Stochastic Bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Decentralized Cooperative Stochastic Bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:34:32.413143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:34:32.413143Z digest=sha256:554ff10315ad7c4012ead922408a7ef9efd3de89b3775deba96ced0b06cda3ab

Observation 49cdaa70-c173-42a8-b78c-6b395fcafe73 · outbound

This paper cites Multi-player bandits–a musical chairs approach.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Multi-player bandits–a musical chairs approach

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.602166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.418683Z digest=sha256:aea8c386828c8f8f9ab4ed3492c62d03c6c7f5edba5acca990293742c55d75af

Observation d14454d1-750a-49ed-9977-7134ddc88cbd · outbound

This paper cites Balanced and incentivized learning with limited shared information in multi-agent multi-armed bandit.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Balanced and incentivized learning with limited shared information in multi-agent multi-armed bandit

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.583749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.424466Z digest=sha256:73349eecc51b80b26c5aba6fe333627380607d3fbfd2b6f148d7061b8547223d

Observation f02470dc-7acb-4bc4-843f-e619ae8530c0 · outbound

This paper cites Multi-Player Multi-Armed Bandits with Finite Shareable Resources Arms: Learning Algorithms & Applications.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Multi-Player Multi-Armed Bandits with Finite Shareable Resources Arms: Learning Algorithms & Applications

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:34:32.496701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.435835Z digest=sha256:f0c50283abe255a00cc6cae98825f1d80c15a19cd1aa4d1c801c40886c4fc326

Observation 8ea434fa-7a54-46c3-bbbb-cf86ae413172 · outbound

This paper cites Decentralized learning for multiplayer multiarmed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Decentralized learning for multiplayer multiarmed bandits

Reference 1979

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.777775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.363727Z digest=sha256:18f16e6cc486b438f58ec13cf0351ed0b81b28762b58aba693df4af0f13f612f

Observation fb96419b-bb2d-4651-8f75-64a72f6de0ff · outbound

This paper cites Bayesian algorithms for decentralized stochas- tic bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Bayesian algorithms for decentralized stochas- tic bandits

Reference 1985

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.737730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.375617Z digest=sha256:6588a9d2ccc4374bb75254829265c1a98f7541a22d771438334e696ec23d1660

Observation ba844463-ca94-48d4-9eff-2e95db99653a · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Finite-time analysis of the multiarmed bandit problem

Reference 1995

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.908689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.310131Z digest=sha256:a07f3cdd854a1ab79dd856fe56a9a63431f7544fbc51d12310dfe7a2456d6da3

Observation c09d8a75-f1d4-42f5-9713-31f83bdecde9 · outbound

This paper cites Communicating with unknown teammates.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Communicating with unknown teammates

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.889492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.315741Z digest=sha256:0a1aa811efda651b821b9b7f10dec417a305a9b1ac282f49402aaec625d69729

Observation 6a4e4c08-f5c3-4fac-83ca-ef26f51c5649 · outbound

This paper cites Decentralized heterogeneous multi- player multi-armed bandits with non-zero rewards on collisions.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Decentralized heterogeneous multi- player multi-armed bandits with non-zero rewards on collisions

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.621530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.408346Z digest=sha256:1691c7414eb35c634c6d951035ca89f6dd7fcf964337ed797d2f49e9078843b5

Observation 8f2f8cbd-1210-40cf-af60-0ef5947bc0e0 · outbound

This paper cites Coordinated versus decentralized exploration in multi- agent multi-armed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Coordinated versus decentralized exploration in multi- agent multi-armed bandits

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.816167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.345834Z digest=sha256:8ef589bd5fe8bc3215dceebfa0dcd528e7335117a7b002c25bcec81a1ed5e403

Observation b08c4370-9e0b-49cf-8dca-188470a98880 · outbound

This paper cites Game of thrones: Fully distributed learning for multi- player bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Game of thrones: Fully distributed learning for multi- player bandits

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.872045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.321499Z digest=sha256:bb12bd9e71410206b4f0274c5b41492e4ae871f5f4ea140c88191fa598c0b6a7

Observation 5b6723f9-ee7f-4168-a16f-7adedc730406 · outbound

This paper cites Multi-agent multi-armed bandits with limited communication.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Multi-agent multi-armed bandits with limited communication

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.945226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.298686Z digest=sha256:96b8bb21ce5223d6f2bb7b88ecd9a8a807a4d5894f96e91d4cfe3a63a5d3c91f

Observation bf4a7aca-b7d9-4bdf-874d-60d88e9f27d5 · outbound

This paper cites Optimal Cooperative Multiplayer Learning Bandits with Noisy Rewards and No Communication.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Optimal Cooperative Multiplayer Learning Bandits with Noisy Rewards and No Communication

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:34:32.352464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:34:32.352464Z digest=sha256:835d7671fcf32614e7bda18808fb8dbedd2f9fc559c38c31c9b37a3b16b108ca

Observation 47774e1c-d879-4c1c-aff5-dd156cc021d3 · outbound

This paper cites Distributed cooperative deci- sion making in multi-agent multi-armed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Distributed cooperative deci- sion making in multi-agent multi-armed bandits

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.680991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.392058Z digest=sha256:d4a5cc1278cb5ab226e5f49602165b1ef8cd8bfdcc636fc35a7918b98f557708

Observation 90557fa2-6e8d-48d5-8601-8cda7787345a · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.836025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.335859Z digest=sha256:64883022838a25172df79a6a02fc3f0a8e9cacb1eb1453fa6fec9a626ce8a3f6

Observation d3d053b0-c22d-45ca-89d6-a455e00d2949 · outbound

This paper cites Distributed learning in multi-armed bandit with multiple players.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Distributed learning in multi-armed bandit with multiple players

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.642473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.403849Z digest=sha256:4e30dd846b507e56d9014dd14930d4fadb3c86fdee8c8419abc62362262780dc

Observation 91536ef5-ece8-4479-ab15-46bfc395934a · outbound

This paper cites Sic-mmab: Synchronisation involves communi- cation in multiplayer multi-armed bandits.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Sic-mmab: Synchronisation involves communi- cation in multiplayer multi-armed bandits

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.854327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.329919Z digest=sha256:e317523351ac0eb4b10f98fbaa578cff6be62c1be3765e61f9dc8e700bc1f1e1

Observation a34af3d0-ca81-4a36-9627-b9c04d7eb125 · outbound

This paper cites Gambling in a rigged casino: The adversarial multi-armed bandit problem.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Gambling in a rigged casino: The adversarial multi-armed bandit problem

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.926264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.304700Z digest=sha256:6fbea4eada1a07c0bc33634bf0bc79ab04f290b81ceae37886540b7c5eee3c1e

Observation cc8474a5-2653-49c6-9c2c-a885495b03b1 · outbound

This paper cites Bandit processes and dy- namic allocation indices.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Bandit processes and dy- namic allocation indices

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.798230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.358412Z digest=sha256:bf8137c3daa2fdec43d212b8427aede8346f8035e269bc65bf148bf55667be42

Observation c3dfcf2b-268f-495b-9187-afdcc8ccd723 · outbound

This paper cites Ad hoc autonomous agent teams: Collaboration without pre-coordination.

Learning to Coordinate Under Threshold Rewards: A Cooperative Multi-Agent Bandit Framework Ad hoc autonomous agent teams: Collaboration without pre-coordination

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:34:32.564075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T19:34:32.430106Z digest=sha256:2c04a92daa0685f4b2a062d3631804e4e120b85724a4726e306d6f53b314e274

Pith citing papers

No inbound Pith citation observations are available.