Pith. sign in

Paper Citation Record · LEDGER

Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2010.03161.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2010.03161 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:01:39.324907Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0fe8e462-4df1-4cd4-be59-fd0cd6758d7a · inbound

Online Learning in MDPs with Partially Adversarial Transitions and Losses cites this paper.

Online Learning in MDPs with Partially Adversarial Transitions and Losses Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T03:01:39.324907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:01:39.324907Z digest=sha256:4abe8557723cfa13a095310fd72bdcf3bc1e10581f6b333b7fe31705f98ff8b8

Observation 252b8257-08f1-4bad-a9a0-fbefa03447a0 · inbound

DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guarantees cites this paper.

DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guarantees Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T08:27:51.715561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:26:03.736064Z digest=sha256:cbd9fd75bf3189ce19913877072227ce9b4e122b95172b06a589c7323222af37

Observation 3af71a45-cb6f-4b64-b56c-62f299048ab7 · inbound

DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guarantees cites this paper.

DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guarantees Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T07:42:30.694419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:39:27.745629Z digest=sha256:206894e5960a5d8d6f971d0431ea231f43b37c1c5cb42ef46a491f25fa450ffc

Observation e0b47cc3-ed42-4455-a96a-56751a515c7f · inbound

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs cites this paper.

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:24:31.610307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T09:24:21.047027Z digest=sha256:aa07cbcc5715a58a48cfa0cf7177640ca44c19413a2c5a19f4a2373a22203ff2

Observation 95a7b86e-fb5b-488b-ada1-444d9e8e0316 · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-31T23:51:57.156793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:51:57.156793Z digest=sha256:1c3e41f50e16bd64597d038ffd0588847c67cff81a23b9eb122c089609acc1b6