Pith. sign in

Paper Citation Record · LEDGER

Learning in Markovian bandits with non-observable states and constrained decision epochs

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2606.27448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.27448 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T01:20:21.181843Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07e44a8e-f684-4c2a-8b50-4baeb1cbc319 · outbound

This paper cites neurips.cc/paper_files/paper/2019/file/88fee0421317424e4469f33a48f50cb0-Paper.pdf.

Learning in Markovian bandits with non-observable states and constrained decision epochs neurips.cc/paper_files/paper/2019/file/88fee0421317424e4469f33a48f50cb0-Paper.pdf

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:297ab7f91d548ba429cc67ef32165e9921ceb5bb124bbccf878116a3537d7cb8

Observation 0270dfa6-9c1c-4e49-8df8-678260715a3c · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Learning in Markovian bandits with non-observable states and constrained decision epochs The regret lower bound for communicating Markov Decision Processes

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.136652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:23b57e59db7b27c86c4022cc12cc32141aa73ad4ef9cbd487500ebcb41b6a280

Observation 767d66c0-aad8-4698-a560-33ac382859f2 · outbound

This paper cites Thompson Sampling in Non-Episodic Restless Bandits.

Learning in Markovian bandits with non-observable states and constrained decision epochs Thompson Sampling in Non-Episodic Restless Bandits

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.139261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:f7808b6f3a119999d295b60e18078b63ebcdd6c91f4a9b776c4ee459a7780371

Observation 82796a65-172f-4a47-b398-e61fd1541736 · outbound

This paper cites Logarithmic weak regret of non-bayesian restless multi-armed bandit.

Learning in Markovian bandits with non-observable states and constrained decision epochs Logarithmic weak regret of non-bayesian restless multi-armed bandit

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:e87ba47e957eda12545587f98c5fe38b5e82be4f361d80d81b0b4089cf4cf18c

Observation 0644c29f-677f-4677-a41f-3abe3660a766 · outbound

This paper cites Learning in a changing world: Restless multiarmed bandit with unknown dynamics.IEEE Transactions on Information Theory, 59(3):1902–1916,.

Learning in Markovian bandits with non-observable states and constrained decision epochs Learning in a changing world: Restless multiarmed bandit with unknown dynamics.IEEE Transactions on Information Theory, 59(3):1902–1916,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:ddde619994dcc72c724825c678db5ba643bfe8b0d6cebb7920b32898d5cea172

Observation 74e3f11d-0540-4342-bec9-7f8b67a8098c · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

Learning in Markovian bandits with non-observable states and constrained decision epochs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.142261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:392c962b5ca5295b40fa5af79183494a96d2c186ec6104647ff9c8990cbd6e51

Observation a7a8a8ea-be33-49cf-90fe-5610276eef58 · outbound

This paper cites As a direct consequence of Lemma A.1, the regret can be approximated by ∑T−1 t=1 (g∗−R(t)), where g∗= maxa∈Aga.

Learning in Markovian bandits with non-observable states and constrained decision epochs As a direct consequence of Lemma A.1, the regret can be approximated by ∑T−1 t=1 (g∗−R(t)), where g∗= maxa∈Aga

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:6602e48f6242f6d96651e5f7285d962ace4303c5a33cee2f8a4b9fffabb41b75

Observation 0c3aa586-26cf-464e-b470-16fd06458b68 · outbound

This paper cites Using a change of measure (Appendix C.1, see Theorems C.1 and C.2), we can bind the behavior of the learning algorithm inM1 to its behavior onM2.

Learning in Markovian bandits with non-observable states and constrained decision epochs Using a change of measure (Appendix C.1, see Theorems C.1 and C.2), we can bind the behavior of the learning algorithm inM1 to its behavior onM2

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:f8be541abc8fa892a84ca9eed9e2de137f9a1fc9514aaea1750aa65b7de99caf

Observation 9f299c58-78d3-45a8-becd-32efe5e06374 · outbound

This paper cites 31 Initial, trueMa Ma Possible alternativeMε a Mε a s∞ ε R= 1 Figure 6: An illustration of the transformationMε a.

Learning in Markovian bandits with non-observable states and constrained decision epochs 31 Initial, trueMa Ma Possible alternativeMε a Mε a s∞ ε R= 1 Figure 6: An illustration of the transformationMε a

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:945d5d975b58422359fd7c0cb2b7b1e478e0671e6b2107c5c58ee2b84289745d

Observation 10f814b7-1563-4d38-bc34-51b131b4e9c5 · outbound

This paper cites an unresolved cited work.

Learning in Markovian bandits with non-observable states and constrained decision epochs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:828248b681906a31058931f575be14a4f0a8c338db66be0d6e8336abc2229dfd

Observation 338e0e30-e0fc-47be-86b7-9481704e84a1 · outbound

This paper cites The dominant termC0 b √ |A|Tlog(T)scales exclusively with the span bound.

Learning in Markovian bandits with non-observable states and constrained decision epochs The dominant termC0 b √ |A|Tlog(T)scales exclusively with the span bound

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:b2d883ac77b35f49c48808ac295074393da1f02d2c68c69877e6d64e75d5a9a5

Observation 1a5851e5-a193-4011-b2c7-7a4270cbdbe0 · outbound

This paper cites an unresolved cited work.

Learning in Markovian bandits with non-observable states and constrained decision epochs Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T01:20:21.181843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:140030ef219d24f0d65d13d0cf9f95bc23cd7b7734686ce275e3294693a4d89f

Pith citing papers

No inbound Pith citation observations are available.