Pith. sign in

Paper Citation Record · LEDGER

Online Learning in MDPs with Partially Adversarial Transitions and Losses

As of 9 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2602.09474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.09474 v2

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:01:40.909364Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dade1b98-4bbb-42b9-9df5-5655a500dcce · outbound

This paper cites We’ll now proveG 4 is true w.p at least 1−δ.

Online Learning in MDPs with Partially Adversarial Transitions and Losses We’ll now proveG 4 is true w.p at least 1−δ

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:01:40.287853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.287853Z digest=sha256:d7379f266e2a5a13a5d0320a88372493fce0b85484601857254ab7f20eb36858

Observation 0ba61a41-0bc5-4473-9d3d-5e6c635a333a · outbound

This paper cites Online learning in episodic markovian decision processes by relative entropy policy search.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Online learning in episodic markovian decision processes by relative entropy policy search

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.064946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.064946Z digest=sha256:085880a6a8e744ba32cd364a2c8251b5993919c4edae0f0b46a100cbe901a9d4

Observation 1eabc990-6bea-46e8-9fcd-411e857fe43b · outbound

This paper cites an unresolved cited work.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.368373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.368373Z digest=sha256:0fd9016968f4926fdf6d7602957e7b687b7abde5971c7c2f3c1b1149c471543a

Observation 694328ff-cfee-4c8e-96fd-774cc1636c45 · outbound

This paper cites an unresolved cited work.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Unresolved cited work

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:01:40.556561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.556561Z digest=sha256:8a8110afe81d55c1bd2422ec45d6c971f6f90473e89ab3f23e2ba52e74431c36

Observation 2d3009bd-3173-443b-945c-3b7115547b65 · outbound

This paper cites qpk,πk h (s, a)− X c∈Ch ˆµk h(s, a, c)ϱk(c) ! ¯ℓk h(s, a) # (tower rule) = X k,h,s,a Ek.

Online Learning in MDPs with Partially Adversarial Transitions and Losses qpk,πk h (s, a)− X c∈Ch ˆµk h(s, a, c)ϱk(c) ! ¯ℓk h(s, a) # (tower rule) = X k,h,s,a Ek

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.751409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.751409Z digest=sha256:6efc5da04bbcd6effe3fd96103c0504123d3176c5e869d0f83d4c95c8bf339dd

Observation b1243624-b570-4126-86c0-c88a71ccca3d · outbound

This paper cites bandit decision step.

Online Learning in MDPs with Partially Adversarial Transitions and Losses bandit decision step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.909364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.909364Z digest=sha256:67ec144454595f584ed0d2878a7e511e8c590ac7de8369b2c5257c1ed2a25c9a

Observation 92a4e4db-fb0b-483f-8397-5d901eba3fcc · outbound

This paper cites That is, a tuple (s, a, s′) for each step inΛ h.

Online Learning in MDPs with Partially Adversarial Transitions and Losses That is, a tuple (s, a, s′) for each step inΛ h

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.202680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.202680Z digest=sha256:74d5578cf6b5e578386e784c3f7f7b2cb6b52134605535abd441fd03d65ca0d8

Observation 29c0e365-0bf4-4273-9e28-0bcf44422b12 · outbound

This paper cites Corruption-robust exploration in episodic reinforcement learning.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Corruption-robust exploration in episodic reinforcement learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.214140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.214140Z digest=sha256:4d494c57e075efdc8c81f0f2ad5f3c7fb5151afa242a769857c317892538b88a

Observation 316ee0d0-ef26-477e-a2f0-174e4c21f9b2 · outbound

This paper cites an unresolved cited work.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.734586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.734586Z digest=sha256:f1b801ce9a83b1e6dfcd2edab5a15f4f200375cb798439a1e41148d198027c16

Observation 0fe8e462-4df1-4cd4-be59-fd0cd6758d7a · outbound

This paper cites Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.324907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.324907Z digest=sha256:4abe8557723cfa13a095310fd72bdcf3bc1e10581f6b333b7fe31705f98ff8b8

Pith citing papers

No inbound Pith citation observations are available.