Pith. sign in

Paper Citation Record · LEDGER

Online Learning in MDPs with Partially Adversarial Transitions and Losses

As of 9 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2602.09474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.09474 v2

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:01:40.909364Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dade1b98-4bbb-42b9-9df5-5655a500dcce · outbound

This paper cites We’ll now proveG 4 is true w.p at least 1−δ.

Online Learning in MDPs with Partially Adversarial Transitions and Losses We’ll now proveG 4 is true w.p at least 1−δ

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:01:40.287853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.287853Z digest=sha256:94e44f64607ac5f5b2cfa592d0b667b43f200ca02088ff4f62cd4aac84a5e3a4

Observation 0ba61a41-0bc5-4473-9d3d-5e6c635a333a · outbound

This paper cites Online learning in episodic markovian decision processes by relative entropy policy search.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Online learning in episodic markovian decision processes by relative entropy policy search

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.064946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.064946Z digest=sha256:759f1643a88b6a823f62fe14d3fbbff630fd9e12fb5b5c8c714aa68ef11af7eb

Observation 1eabc990-6bea-46e8-9fcd-411e857fe43b · outbound

This paper cites an unresolved cited work.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.368373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.368373Z digest=sha256:032afcef57f56950f1b9bd8f9fe1a645679c6789162489f3e429abc95a28198a

Observation 694328ff-cfee-4c8e-96fd-774cc1636c45 · outbound

This paper cites an unresolved cited work.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Unresolved cited work

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-03T03:01:40.556561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.556561Z digest=sha256:27302948b21f15669c65e8fa89ac2780065254b9e1b97a5b71dc5190a09ec9a6

Observation 2d3009bd-3173-443b-945c-3b7115547b65 · outbound

This paper cites qpk,πk h (s, a)− X c∈Ch ˆµk h(s, a, c)ϱk(c) ! ¯ℓk h(s, a) # (tower rule) = X k,h,s,a Ek.

Online Learning in MDPs with Partially Adversarial Transitions and Losses qpk,πk h (s, a)− X c∈Ch ˆµk h(s, a, c)ϱk(c) ! ¯ℓk h(s, a) # (tower rule) = X k,h,s,a Ek

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.751409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.751409Z digest=sha256:97d1432872a43a7b3837b648fbac5aec510fd1754ecf649fff68e5ec2f614369

Observation b1243624-b570-4126-86c0-c88a71ccca3d · outbound

This paper cites bandit decision step.

Online Learning in MDPs with Partially Adversarial Transitions and Losses bandit decision step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.909364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.909364Z digest=sha256:db2a779ea1383adebd10d20a26e57d912a466bb8ff013c75a39936b4732fdea7

Observation 92a4e4db-fb0b-483f-8397-5d901eba3fcc · outbound

This paper cites That is, a tuple (s, a, s′) for each step inΛ h.

Online Learning in MDPs with Partially Adversarial Transitions and Losses That is, a tuple (s, a, s′) for each step inΛ h

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:40.202680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:40.202680Z digest=sha256:64bd26ef231b8e80950bf9cc42de4404cc773f45fba7dd1c72604d6c568dc261

Observation 29c0e365-0bf4-4273-9e28-0bcf44422b12 · outbound

This paper cites Corruption-robust exploration in episodic reinforcement learning.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Corruption-robust exploration in episodic reinforcement learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.214140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.214140Z digest=sha256:396d24c3999298438590ceaba805abba376fd59fe4b9dcc7e6410329052b7d15

Observation 316ee0d0-ef26-477e-a2f0-174e4c21f9b2 · outbound

This paper cites an unresolved cited work.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.734586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.734586Z digest=sha256:2f2fec7c54616ba94a49cbd36671d58f692e0f67d7501c7e6badcb4438948d4b

Observation 0fe8e462-4df1-4cd4-be59-fd0cd6758d7a · outbound

This paper cites Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.324907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.324907Z digest=sha256:33c1706c95047dde7377192f7ee5789b5968a5e4700a5cafdd2996fd519f900e

Pith citing papers

No inbound Pith citation observations are available.