Pith. sign in

Paper Citation Record · LEDGER

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2505.21841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21841 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:51.668145Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 505562c8-8319-48b4-a76b-f495dc7b1a99 · outbound

This paper cites write newline.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:28.970873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:34:28.970873Z digest=sha256:a9a64ee47a285b8afc3bffb9f32d625605b601e5a5177306b841ad3285dbc35f

Observation 630fe61f-14f5-4025-aa51-0741d651fe4b · outbound

This paper cites Constrained policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained policy optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.539697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.027164Z digest=sha256:2787b1ab4cc3ffa1d9a6720b58418045e6ac4a11d52966d89458a9456f82e616

Observation 00939444-31d3-4e08-bcf8-9f10515a817e · outbound

This paper cites Constrained Markov decision processes, volume 7.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.227495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.159068Z digest=sha256:f87940513bbf719b18b4150f0906f24b2fe80fcf3eced9945069df131b809bad

Observation 702aad4e-aa41-47b2-bc61-d37059c9958e · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:58.051680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.450316Z digest=sha256:97c09546b942cc9d7d85441826d16c1ff4dd848091de4fab73e5585201de42eb

Observation f732cb5a-352a-4cee-a1b6-46c437e7cb63 · outbound

This paper cites G., Osband, I., and Munos, R.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints G., Osband, I., and Munos, R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:47.683128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:47.683128Z digest=sha256:5add85cee4e6f223324ffc6979204ff0a018b196eb3c6c61363f1073e267e40c

Observation 11fef369-1567-4313-a78b-c6b810038091 · outbound

This paper cites S., Agarwal, M., Koppel, A., and Aggarwal, V.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints S., Agarwal, M., Koppel, A., and Aggarwal, V

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.781747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.767912Z digest=sha256:7b64a84b973d8b6494ed748da86f6621890141e3e9aa50039313c98f80fc789c

Observation eff5fcfe-0b06-4347-bcf2-9688eb8b64b2 · outbound

This paper cites DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.527372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:47.885604Z digest=sha256:ab2ab8e731ff46a6817bc07bb58808317ce5055fc1cecb5854364f9f3eb3d8a7

Observation 3347a759-18a2-4458-a690-11666c403118 · outbound

This paper cites Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.355079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.002331Z digest=sha256:f553eb9dec6f97c1367c6c61d228e602f6627a6e0ebff8d479268796b129d230

Observation 900a97dd-cfee-456f-953d-bbc0c71787bc · outbound

This paper cites Learning infinite-horizon average-reward markov decision process with constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning infinite-horizon average-reward markov decision process with constraints

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.621924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.098233Z digest=sha256:d27a1f0ff7ac87b6e02a3b08f562d7aa87291c03cb4c363b2b028b2c2fc1179e

Observation 06e5816d-1349-4dd4-b8ec-d67ed52b0e08 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Risk-constrained reinforcement learning with percentile risk criteria

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.335859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.234339Z digest=sha256:3a97efbb9cbcd23e3556522135276cd74fb2b367db33fc7aa79c8ad8cce696f2

Observation 8fcd96a9-0e4c-43af-b02f-be102f098bb2 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:57.018850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.336113Z digest=sha256:ca0c020f035ea16e2c1c14bae12b78c4a4c33f18d9db35df95f3346979d66414

Observation cdc0c94a-c330-4168-81ad-7bdc10a34556 · outbound

This paper cites Provably efficient safe exploration via primal-dual policy optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient safe exploration via primal-dual policy optimization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.831806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.467808Z digest=sha256:772baf3bf716aaa589eb0831aea802a77c78c2153dc3f0aab95edbfc9b4d5151

Observation 6e50368c-c30d-4027-b003-6cdfbcd64dc7 · outbound

This paper cites Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably Efficient Primal-Dual Reinforcement Learning for CMDPs with Non-stationary Objectives and Constraints

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:53.098564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.569789Z digest=sha256:6f9ce8e885999d317dff920d685f63aff1a85bee71003e99ced65f13cd7da22c

Observation 9b1fab1b-2b67-43ce-a413-b9e645253a4c · outbound

This paper cites Exploration-Exploitation in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Exploration-Exploitation in Constrained MDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:48.684747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:48.684747Z digest=sha256:520e519a57fb6b97daa625fd64707204d8e97b15f80c30414b90ab1c547b6972

Observation 3ad0cfde-e1bf-4a66-b104-c999d0353657 · outbound

This paper cites A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:52.838965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.792791Z digest=sha256:b728ef48c1893124e011c944f2c3647640489d2206bcba96dfee9eada0790ea6

Observation e064fe05-10de-4988-a9a1-0ebb44156973 · outbound

This paper cites Provably efficient model-free constrained rl with linear function approximation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free constrained rl with linear function approximation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.615244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:48.894556Z digest=sha256:481e212cb371c75523a19a8890ce764da74b0e225e84b9cb19adad6b0548ecbb

Observation d06a5db2-2c72-4b9a-bcbc-5682e2b3ec77 · outbound

This paper cites Online convex optimization with hard constraints: Towards the best of two worlds and beyond.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Online convex optimization with hard constraints: Towards the best of two worlds and beyond

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.433971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.064830Z digest=sha256:55722eae459aa4b653744c7a76e3751cc320f6e0e5028aefdfaf869b1ae51c46

Observation 55b6921b-dfe2-4b24-92c6-c7e51f6ea4aa · outbound

This paper cites Safe reinforcement learning on autonomous vehicles.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Safe reinforcement learning on autonomous vehicles

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.252065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.160323Z digest=sha256:a897495d755652acbb6afb4ce6ec0ae691cd771fc6be2a84545a5cefb95d8bea

Observation cd4ebc55-86eb-4332-9128-e09e759b68b5 · outbound

This paper cites Learning adversarial markov decision processes with bandit feedback and unknown transition.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning adversarial markov decision processes with bandit feedback and unknown transition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:56.073085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.233974Z digest=sha256:581acee5b36f17479cef3502f68b06e0e1b66bbfd9ba846f2a0d1302944d28b5

Observation 4cc7faf8-aad4-4a61-8a87-fc00818c6150 · outbound

This paper cites Asymptotically Optimal Information-Directed Sampling.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Asymptotically Optimal Information-Directed Sampling

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.598973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.315365Z digest=sha256:7f04394b7c8822fef62525e6dcc79862b00bdf35af067bf3037b840f34d1d2c3

Observation 3152b411-b341-40c6-8c39-757ffd025609 · outbound

This paper cites A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.425405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.425405Z digest=sha256:50dc8007aa5e2039f9b62df5a11af65e61ffd48f7270298e2b4bc93a180422fc

Observation a8af659f-5cf7-463d-b447-43f7d5501947 · outbound

This paper cites An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints An Optimistic Algorithm for Online Convex Optimization with Adversarial Constraints

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.517407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.517407Z digest=sha256:84bff569f99caf89902648a51ec5af0bc1fc86675566197fac7322b11e8fa1ae

Observation 0ae5c9b7-2a55-4a3d-8f5c-2bddaf7c3ce1 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained MDPs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.912890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.610537Z digest=sha256:5151aa96d1338927f641fa82aafd966cfa89617cbc4ddce332e9bc9088e9b464

Observation 43ca9dfd-247a-45e0-8f8d-a91c9e1a2c08 · outbound

This paper cites Learning policies with zero or bounded constraint violation for constrained mdps.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning policies with zero or bounded constraint violation for constrained mdps

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.718689Z digest=sha256:5f9836b514e12b0ada3329986da75859e1f14b912740fc9d5c7adc8ee1722a93

Observation 30023b5d-0b14-481a-8259-27aae1a15833 · outbound

This paper cites Policy optimization in adversarial mdps: Improved exploration via dilated bonuses.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Policy optimization in adversarial mdps: Improved exploration via dilated bonuses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.501758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:49.844197Z digest=sha256:511a644aa343cdc4c5a0b47d02d653d653461e3d59dec932115efd88f3568c4a

Observation b0a2135b-3cbe-424a-981b-d9ac16ab7d0a · outbound

This paper cites Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Cancellation-Free Regret Bounds for Lagrangian Approaches in Constrained Markov Decision Processes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:49.965851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:49.965851Z digest=sha256:439a7391a00b564860db0b585d649a326e9b106131531f19bc71c4bf716ce76b

Observation 132bc39d-c3d1-47da-bf21-fd3bedf6e30a · outbound

This paper cites Truly No-Regret Learning in Constrained MDPs.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Truly No-Regret Learning in Constrained MDPs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:50.106454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:50.106454Z digest=sha256:2850025b8f6987201bee9356e0b21301eb5e9459f14bbbcfd49df1e189dd4431

Observation 4bd08f26-9eb0-45c3-84b4-49c23f9636e1 · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:55.253773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.236838Z digest=sha256:0b2f0789816f782811d38162755b87049cbd8d8eb44eebf0889629e4c4322bb8

Observation d496fcf0-1dd2-49f9-b471-fda868e43a9b · outbound

This paper cites Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:55.038809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.374739Z digest=sha256:1cce3594b73efb56cdb942553e977b2527cbe532b82452c4cb2d457ee75b91d4

Observation 1d7336e1-2e04-4a89-9b5c-f897fe1b595d · outbound

This paper cites and Sridharan, K.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Sridharan, K

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.836763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.535752Z digest=sha256:07dad60d6a0b759af8afaffa3b9620034b6313115e7905ddd539a15e9370bbaf

Observation 5435887d-aee8-47e5-be4a-1626685ea8d8 · outbound

This paper cites Learning in Markov Decision Processes under Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning in Markov Decision Processes under Constraints

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.311657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.654786Z digest=sha256:713502ec38e117470db01e1ed081ac7d463b7026c3ba9ef19946af9766980eae

Observation 69f3302e-3541-4a29-b121-6ab7b2a4be2f · outbound

This paper cites Optimal Algorithms for Online Convex Optimization with Adversarial Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Algorithms for Online Convex Optimization with Adversarial Constraints

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.145545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:50.768281Z digest=sha256:2090f81d3688c98a02b8080503c1d2cc261e4185fc6a369819bcc8bf5c76d774

Observation a182d344-fcc4-42ee-bf95-7cbc40d03d95 · outbound

This paper cites Learning Adversarial MDPs with Stochastic Hard Constraints.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Learning Adversarial MDPs with Stochastic Hard Constraints

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.044652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:51.044652Z digest=sha256:fecad7404555c81101ba28ac143423da7ce06eae91780970ae9ceb4d43839557

Observation b8adca20-6eb5-4929-a697-dde12424916b · outbound

This paper cites Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:35:51.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.152297Z digest=sha256:acec31e040537a638f854b6a0d8f6bf5678a5a4862edbc0a4b5cfe80a18153cc

Observation 18847069-32b1-4364-aee1-77751cd6bdd9 · outbound

This paper cites Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Triple-Q: a model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.549056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.294799Z digest=sha256:a33d42bc33b651c383634af868faa5741d86c022b21ba8464238f2b3bd92a9cb

Observation ac158d50-b7f6-4f0d-928e-2102d5c876fb · outbound

This paper cites A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints A provably-efficient model-free algorithm for infinite-horizon average-reward constrained markov decision processes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.348961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.366080Z digest=sha256:b384e6a2cd1b5c3b9210e91e86bb00648e771eed0afd2b8ed1c1af8c70dc1306

Observation 10b44870-fe67-4b09-aa82-f812732727e1 · outbound

This paper cites Provably efficient model-free algorithms for non-stationary CMDP s.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Provably efficient model-free algorithms for non-stationary CMDP s

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:54.171541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.499356Z digest=sha256:2712220fd10fd1a2dbf5e1e299eb749ee94ea58fc23122405aee7c45cfd423fc

Observation 531160d9-5088-45ee-a051-a9a4dcc157ab · outbound

This paper cites an unresolved cited work.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:35:53.960811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.570309Z digest=sha256:5d6cefba715981820efc4ff3cad7a9b4544313273d9e6df92d161d7ad53346da

Observation 42c9f1b7-449e-4905-b842-e43a45cd2e29 · outbound

This paper cites and Ugot, O.-A.

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints and Ugot, O.-A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:35:53.728331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:35:51.668145Z digest=sha256:cb6942b63228178f91479659a8abff2b81afb97fb050091635599bf2c833bb92

Pith citing papers

No inbound Pith citation observations are available.