Pith. sign in

Paper Citation Record · LEDGER

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

As of 13 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2505.21841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21841 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:51.668145Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 505562c8-8319-48b4-a76b-f495dc7b1a99 · outbound

This paper cites write newline.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:28.970873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:34:28.970873Z digest=sha256:8f8a24d8817a37938b3e13a55b49364ddc81357a49e75b1bdd2220cf74355ef9

Observation 630fe61f-14f5-4025-aa51-0741d651fe4b · outbound

This paper cites Constrained policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained policy optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.539697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.027164Z digest=sha256:b0c5cc627b098e948e41ef6bb32fae464470048aea05101a7e18f94cb37732af

Observation 00939444-31d3-4e08-bcf8-9f10515a817e · outbound

This paper cites Constrained Markov decision processes, volume 7.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.227495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.159068Z digest=sha256:8807e576257fd9ce7e994219c33e73a953b1c2ff8bceb4cc6df350f57b7db2ed

Observation 702aad4e-aa41-47b2-bc61-d37059c9958e · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.051680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.450316Z digest=sha256:111f34aa9755fcb4e3cfaf99e631340a33af44401338219f743efe3431a1a9e8

Observation f732cb5a-352a-4cee-a1b6-46c437e7cb63 · outbound

This paper cites G., Osband, I., and Munos, R.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints G., Osband, I., and Munos, R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:47.683128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:47.683128Z digest=sha256:f31bee78004abdd9377c4486ac12f377e5f73f17ec05ae9d0b31b4f523b6970a

Observation 11fef369-1567-4313-a78b-c6b810038091 · outbound

This paper cites S., Agarwal, M., Koppel, A., and Aggarwal, V.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints S., Agarwal, M., Koppel, A., and Aggarwal, V

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.781747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.767912Z digest=sha256:b288139342096377d6f7906005bf194132c901e5c74394b1e8ff924fc3f9c5a2

Observation eff5fcfe-0b06-4347-bcf2-9688eb8b64b2 · outbound

This paper cites DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.527372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.885604Z digest=sha256:f9707cbc605575faa8a83969454f583624ed1e896e9ab6e3eff5514dd6a36b6f

Observation 3347a759-18a2-4458-a690-11666c403118 · outbound

This paper cites Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.355079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.002331Z digest=sha256:331f4e8c4effc73a382d2ec6d933d3ef7020b86f5a6d647f2f2b61490a2a08e7

Observation 900a97dd-cfee-456f-953d-bbc0c71787bc · outbound

This paper cites Learning infinite-horizon average-reward markov decision process with constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning infinite-horizon average-reward markov decision process with constraints

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.621924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.098233Z digest=sha256:b2c67097f2193e4a5c48dc9ec92fd2fa4dbd2f0954e076eedf80ba803b01390b

Observation 06e5816d-1349-4dd4-b8ec-d67ed52b0e08 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Risk-constrained reinforcement learning with percentile risk criteria

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.335859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.234339Z digest=sha256:7b592c0b19f2d22813fb269ec0b39c2ede16548705417066b44229ab8d9ee883

Observation 8fcd96a9-0e4c-43af-b02f-be102f098bb2 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.018850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.336113Z digest=sha256:2b70d20e15dd2d27b63a93978c88ac5458f3b1bfb0d57412302f78e0e544f310

Observation cdc0c94a-c330-4168-81ad-7bdc10a34556 · outbound

This paper cites Provably efficient safe exploration via primal-dual policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient safe exploration via primal-dual policy optimization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.831806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.467808Z digest=sha256:73355ed6e445257b12d6e79c97d96ecf4055811684b2ec6ce554d8dc293a1d5e

Observation 6e50368c-c30d-4027-b003-6cdfbcd64dc7 · outbound

This paper cites Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.098564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.569789Z digest=sha256:d4144815f62c802bca2834a9d3609661e7077c761b0038d7470d555807e76e36

Observation 9b1fab1b-2b67-43ce-a413-b9e645253a4c · outbound

This paper cites Exploration-Exploitation in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Exploration-Exploitation in Constrained MDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:48.684747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:48.684747Z digest=sha256:b5440d0c45296680cbe9f8810c0d8d2216ac94a482adfd7c037b2b3046c8706d

Observation 3ad0cfde-e1bf-4a66-b104-c999d0353657 · outbound

This paper cites A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:52.838965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.792791Z digest=sha256:a75083b232433fd32114bd09a152e03c0da7a9d89169bdb1ac179a5e15c00bfd

Observation e064fe05-10de-4988-a9a1-0ebb44156973 · outbound

This paper cites Provably efficient model-free constrained rl with linear function approximation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free constrained rl with linear function approximation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.615244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.894556Z digest=sha256:227a3acacd8d4a8ca4b2c6c4de26495221698985d563e847a8d2c96d68e7a07d

Observation d06a5db2-2c72-4b9a-bcbc-5682e2b3ec77 · outbound

This paper cites Online convex optimization with hard constraints: Towards the best of two worlds and beyond.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Online convex optimization with hard constraints: Towards the best of two worlds and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.433971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.064830Z digest=sha256:02903696b9315635a0ae65bd34a040959e0dd35a083a0f64c4d374e995bbce70

Observation 55b6921b-dfe2-4b24-92c6-c7e51f6ea4aa · outbound

This paper cites Safe reinforcement learning on autonomous vehicles.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Safe reinforcement learning on autonomous vehicles

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.252065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.160323Z digest=sha256:9bfd6f43ea5dd2bd81b9e2fd6e8df6ae3e77158d56dc09da49812da314f47fbd

Observation cd4ebc55-86eb-4332-9128-e09e759b68b5 · outbound

This paper cites Learning adversarial markov decision processes with bandit feedback and unknown transition.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning adversarial markov decision processes with bandit feedback and unknown transition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.073085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.233974Z digest=sha256:7722b2fa48de91a1f7cc4857c2d12ae1ca889dcbf6aa7652296cb947d1d167a0

Observation 4cc7faf8-aad4-4a61-8a87-fc00818c6150 · outbound

This paper cites Asymptotically Optimal Information-Directed Sampling.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Asymptotically Optimal Information-Directed Sampling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.598973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.315365Z digest=sha256:4d59493d75ef2a1ac178666c4fc7ef0dd2679f0be794ce47e47730e157ccc182

Observation 3152b411-b341-40c6-8c39-757ffd025609 · outbound

This paper cites A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.425405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.425405Z digest=sha256:ce6bbad65e503a5c256af8bb5734259c00050e17d50a1d162462d52f53db2b58

Observation a8af659f-5cf7-463d-b447-43f7d5501947 · outbound

This paper cites An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.517407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.517407Z digest=sha256:241a8bf44ad700257fa293e9a438f0d6c41916531c1d623bba667e6109606a6d

Observation 0ae5c9b7-2a55-4a3d-8f5c-2bddaf7c3ce1 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained MDPs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.912890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.610537Z digest=sha256:dc087135bd20de3858efdbcc8b549dca80a108477b420cf2daef0ef32146597c

Observation 43ca9dfd-247a-45e0-8f8d-a91c9e1a2c08 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained mdps.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained mdps

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.718689Z digest=sha256:2cd62d6c9758b7ac329f6ebe22ab58430314a9acc416ef03a7d1316123bc42ae

Observation 30023b5d-0b14-481a-8259-27aae1a15833 · outbound

This paper cites Policy optimization in adversarial mdps: Improved exploration via dilated bonuses.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Policy optimization in adversarial mdps: Improved exploration via dilated bonuses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.501758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.844197Z digest=sha256:995383fea740d5e0545d73216e083aa5e9400914b70e06fd98cecc53f7a6d7f5

Observation b0a2135b-3cbe-424a-981b-d9ac16ab7d0a · outbound

This paper cites Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.965851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.965851Z digest=sha256:d6e52dd275ea3d4a833c7677b1731b1db2d3b2c88c03606b3ccefea082d95c09

Observation 132bc39d-c3d1-47da-bf21-fd3bedf6e30a · outbound

This paper cites Truly No-Regret Learning in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Truly No-Regret Learning in Constrained MDPs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.106454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:50.106454Z digest=sha256:c9700d3d08e211b996473efac3a113f1cc9be9d10b1003903f856d67f4e55d7a

Observation 4bd08f26-9eb0-45c3-84b4-49c23f9636e1 · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:55.253773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.236838Z digest=sha256:80ba8edf516d8036e73cb1a6b269dc80960f300e11aff51017944f225b1fe46e

Observation d496fcf0-1dd2-49f9-b471-fda868e43a9b · outbound

This paper cites Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.038809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.374739Z digest=sha256:626d433b49b9d342f83ffe139b4937c1f82ec77877218ad9b5685bb1f777e535

Observation 1d7336e1-2e04-4a89-9b5c-f897fe1b595d · outbound

This paper cites and Sridharan, K.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Sridharan, K

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.836763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.535752Z digest=sha256:a65c9cfd8af49fa78a518e412fefbfaade3a411f60818872142e57a475176c13

Observation 5435887d-aee8-47e5-be4a-1626685ea8d8 · outbound

This paper cites Learning in Markov Decision Processes under Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning in Markov Decision Processes under Constraints

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.311657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.654786Z digest=sha256:3c5ce5101b95a22efee08c3992169d24e361df92fe9610cef5fc204cee9d67d7

Observation 69f3302e-3541-4a29-b121-6ab7b2a4be2f · outbound

This paper cites Optimal Algorithms for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Algorithms for Online Convex Optimization with Adversarial Constraints

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.145545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.768281Z digest=sha256:cb18706efdd6d264cec6f18ca307153f65013a2a75846ca6f0c5c2b29fff3fe4

Observation a182d344-fcc4-42ee-bf95-7cbc40d03d95 · outbound

This paper cites Learning Adversarial MDPs with Stochastic Hard Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning Adversarial MDPs with Stochastic Hard Constraints

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.044652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:51.044652Z digest=sha256:d3b63f007e52ed3c1a0a0efbb7b74134154be4d90d2deb03b66e29536d8ef02a

Observation b8adca20-6eb5-4929-a697-dde12424916b · outbound

This paper cites Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:51.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.152297Z digest=sha256:9dd7b4678bf6d5c2fd76ecf4db746a9c739ec170ec3bee980931e6091b4ba0aa

Observation 18847069-32b1-4364-aee1-77751cd6bdd9 · outbound

This paper cites Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.549056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.294799Z digest=sha256:c3be6eb9bdd02e184249adc76d78d5d9a9b828d57868d22129fd26828364fbb2

Observation ac158d50-b7f6-4f0d-928e-2102d5c876fb · outbound

This paper cites A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.348961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.366080Z digest=sha256:0525eda57b86f531763bd42f47283d59110ebb976ec49296fc3e8263df7e6ec5

Observation 10b44870-fe67-4b09-aa82-f812732727e1 · outbound

This paper cites Provably efficient model-free algorithms for non-stationary CMDP s.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free algorithms for non-stationary CMDP s

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.171541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.499356Z digest=sha256:4e254faa619dfed7d1920d0075a3bf851c52566be0b559544c346fc8d9bf917d

Observation 531160d9-5088-45ee-a051-a9a4dcc157ab · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:53.960811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.570309Z digest=sha256:d0ed63d228f68651efe46b2ed8f793cbb98934a44f0e00778983522a7a4dece3

Observation 42c9f1b7-449e-4905-b842-e43a45cd2e29 · outbound

This paper cites and Ugot, O.-A.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Ugot, O.-A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:53.728331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.668145Z digest=sha256:e25ee2489c3ac6e6f1903dd370bec66456b072abfb62c4d24b65972ca39ce4a4

Pith citing papers

No inbound Pith citation observations are available.