Pith. sign in

Paper Citation Record · LEDGER

Distributionally Robust Deep Q-Learning

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 2 inbound Pith citation observations for arXiv:2505.19058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19058 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:26:53.058645Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T09:56:20.351339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:57:00.793817Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact5
  • verified fuzzy43
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d56f98fb-9fae-4388-a4fb-338b43c8f91d · outbound

This paper cites Investigating the parameters of the beta distribution.

Distributionally Robust Deep Q-Learning Investigating the parameters of the beta distribution

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.823780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.148084Z digest=sha256:a9272f3aaf6ff23a1b9be14d63f9149225f30d4cb2e501dbd8b19d198ba1197b

Observation 9ac03d49-301b-462e-aa0e-f29bd48cd45e · outbound

This paper cites Infinite dimensional analysis: a hitchhiker’s guide.

Distributionally Robust Deep Q-Learning Infinite dimensional analysis: a hitchhiker’s guide

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.810211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.212428Z digest=sha256:f06edf3fd8715fa52942118142512bfc1f9ce05ef46d88f2966cf966a21b54d9

Observation b7f522da-e813-468e-a6b7-fbdbc4005c9e · outbound

This paper cites Computational aspects of robust optimized certainty equiv- alents and option pricing.

Distributionally Robust Deep Q-Learning Computational aspects of robust optimized certainty equiv- alents and option pricing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.797222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.343157Z digest=sha256:edc017a50656f15ca1dff0ae499a5ad0a706ea2d049edc1051381801955fcb73

Observation c17de44c-bf89-4428-b2f5-f41e89f22035 · outbound

This paper cites On the theory of dynamic programming.

Distributionally Robust Deep Q-Learning On the theory of dynamic programming

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.783436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.437965Z digest=sha256:1cc0ba59e49527c48339a97946d1791acd6a246823a5b02f7168807bc32a0f0f

Observation ae0c3ccc-57a8-4ac0-b321-2fa2e0c1c04d · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Distributionally Robust Deep Q-Learning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:47.529254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:47.529254Z digest=sha256:ca7a99328e0df02681e17550e1bab4e84f53baf508bf61ccc9947d1f5505ecb3

Observation fc8655b3-0994-46f6-b078-75faa5cf58a5 · outbound

This paper cites Convex optimization.

Distributionally Robust Deep Q-Learning Convex optimization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.768872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.603514Z digest=sha256:51fa5c2de371810d42f7682dcb1cf1eda660f2e34cd3dd7aa4b05ee639937e2a

Observation 335e8b6a-d2cc-430c-a794-e6c70c8e712c · outbound

This paper cites Distributionally robust Markov decision processes and their connection to risk measures.

Distributionally Robust Deep Q-Learning Distributionally robust Markov decision processes and their connection to risk measures

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.756832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.723448Z digest=sha256:a41b04e9eb31cbe89f5262c529fbff92c09b3fafe8676fd6e304516038ab8b11

Observation 1461b9f8-e872-4b1c-8a58-f98767d84ecf · outbound

This paper cites an unresolved cited work.

Distributionally Robust Deep Q-Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:26:59.745726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.813404Z digest=sha256:048a69133cda8ef5bfe1b92a902a09992cb3de584f936646f8dbab36b3d7a1fc

Observation c8bb8b23-9445-48d0-bd0c-0f859f07c0a6 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport, 2013.

Distributionally Robust Deep Q-Learning Sinkhorn distances: Lightspeed computation of optimal transport, 2013

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.730603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.913194Z digest=sha256:6e5471774759144b7bbf0b818d33f7f7668bd6bce54b07284933ad98ba841374

Observation 836e6293-d9a6-4848-a7b4-c4090a822da9 · outbound

This paper cites Robust Q-learning for finite ambiguity sets.

Distributionally Robust Deep Q-Learning Robust Q-learning for finite ambiguity sets

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:26:54.193985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:47.992530Z digest=sha256:508622e8b9ccad6b4e095d5f9f6a1c0f91149a8fc7494f0e9efe6294194c1963

Observation 757a9a1a-eb1b-4fc5-b04e-5637ca04f287 · outbound

This paper cites Twice Regularized Markov Decision Processes: The Equivalence between Robustness and Regularization.

Distributionally Robust Deep Q-Learning Twice Regularized Markov Decision Processes: The Equivalence between Robustness and Regularization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.921863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.113813Z digest=sha256:9e202b48fd66f617cab121e7db93cb07b0da5590f06048993709644a2bb36037

Observation 016885b5-6336-4a61-83bc-6130eab112a1 · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

Distributionally Robust Deep Q-Learning Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:48.211513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:48.211513Z digest=sha256:8644e0a542f3658094d0454285803373f434fcefdeda3801df2b0135035515df

Observation 6ebf42a6-9e34-4417-be0d-d482898157cd · outbound

This paper cites A theoretical analysis of deep Q-learning.

Distributionally Robust Deep Q-Learning A theoretical analysis of deep Q-learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.717657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.312156Z digest=sha256:49fc91d0a0e5b643ad72cb307283d60e5d3661f93b887477c38aad2f670a3961

Observation 383dc3d1-25d8-4bb0-b7dc-22b3084143b2 · outbound

This paper cites In- terpolating between optimal transport and mmd using Sinkhorn divergences.

Distributionally Robust Deep Q-Learning In- terpolating between optimal transport and mmd using Sinkhorn divergences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.703610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.411547Z digest=sha256:be32deb96b43bf57e29bd9f35a0b5249ed03b08db67008db803b6527023fea50

Observation 742e4c8e-f8e8-491c-8a58-6d57f19be4c7 · outbound

This paper cites Sample complexity of Sinkhorn divergences.

Distributionally Robust Deep Q-Learning Sample complexity of Sinkhorn divergences

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.689706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.498761Z digest=sha256:2eb617f4a461c316ecaba36e24d0097de62ffe82516498caf9cf935d75b7fc24

Observation 61f552e5-ee19-46ee-931f-1072e7e14675 · outbound

This paper cites Stability of entropic optimal transport and Schr¨ odinger bridges.

Distributionally Robust Deep Q-Learning Stability of entropic optimal transport and Schr¨ odinger bridges

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.675743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.596240Z digest=sha256:52cd635b6ab5ce7c5e62588c42431c36ed107e746f2c2818cb502b6a940d1416

Observation 80cc6a39-69e5-4015-b210-74594bae82a9 · outbound

This paper cites Robust Markov decision processes: Beyond rectangularity.Mathematics of Operations Research, 48(1):203–226, 2023.

Distributionally Robust Deep Q-Learning Robust Markov decision processes: Beyond rectangularity.Mathematics of Operations Research, 48(1):203–226, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.663234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.650603Z digest=sha256:9b8794bd5fea48d48d320bc528e63843672c516a30fb0c9882f13bf71873d3a8

Observation f6bfc4ed-785b-4e05-9a94-c43e1aedc790 · outbound

This paper cites Double Q-learning.

Distributionally Robust Deep Q-Learning Double Q-learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.650585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.760252Z digest=sha256:395a665abf2c832d658e024304f6a831eaffeb11267f712b6d70d2050bccee19

Observation 14cce011-4b45-4032-9112-e2dbd88ecc1b · outbound

This paper cites Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks.

Distributionally Robust Deep Q-Learning Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:48.845231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:48.845231Z digest=sha256:a4a5754bec3d827a1deeebfd9efffd81cbd4db35b7fd334bcc266f5e03eadf6f

Observation b10a21ee-df77-4096-b7eb-947881c8ce83 · outbound

This paper cites Learning to utilize shaping rewards: A new approach of reward shaping.

Distributionally Robust Deep Q-Learning Learning to utilize shaping rewards: A new approach of reward shaping

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:48.908999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:48.908999Z digest=sha256:2a0fd8bceafb7b00fecd1ffd6fe5843ec316f37aecd5adf6e456f894efccaee8

Observation c468e487-5adc-4025-a1aa-9d4ca6407773 · outbound

This paper cites Robust dynamic programming.

Distributionally Robust Deep Q-Learning Robust dynamic programming

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.619185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:48.942188Z digest=sha256:da22a131d4e0fc95a10857ee911f0055558ade18034d65305ae58c4b4fcb4c7d

Observation 54768045-e007-48f4-b640-1e1cf17b974a · outbound

This paper cites Probability essentials.

Distributionally Robust Deep Q-Learning Probability essentials

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.009637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.009637Z digest=sha256:87435e085ca997c79d2de93eb87eb4f83955e4f772853f0c0083ea459c7a01e3

Observation 44bda2c0-d2e0-465a-95df-957062227227 · outbound

This paper cites Universal approximation with deep narrow networks.

Distributionally Robust Deep Q-Learning Universal approximation with deep narrow networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.077557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.077557Z digest=sha256:677ab4785fba6b5e31c569f1a006a3ad6abaeaf026ecd30c36f78a2adfcc728d

Observation 5daadba4-6eb4-4020-a685-1df0d1095287 · outbound

This paper cites Probability theory: a comprehensive course.

Distributionally Robust Deep Q-Learning Probability theory: a comprehensive course

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.586840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:49.194831Z digest=sha256:36955ed83d6c2489e56a8a9e93da9e82be9191b9e5d1f18c617a1d53dcca683a

Observation 2196972d-d2b4-4108-97d5-1b9403556f1d · outbound

This paper cites An Efficient Solution to s-Rectangular Robust Markov Decision Processes.

Distributionally Robust Deep Q-Learning An Efficient Solution to s-Rectangular Robust Markov Decision Processes

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.737945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:49.284674Z digest=sha256:ef84cd37dd0f3e3eccc985c39f5e7b20017d8f677c817f3730a1dc6bd904e35b

Observation 7514de8d-3eea-48f6-971f-ef12ca6fd1ec · outbound

This paper cites Playing fps games with deep reinforcement learning.

Distributionally Robust Deep Q-Learning Playing fps games with deep reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.570408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:49.444931Z digest=sha256:849019925221695ba5167311a9c06a62259b01020a66a56aadf82880318d2c80

Observation eb124f69-56be-46e8-bf84-c9cdfc65cddb · outbound

This paper cites Policy gradient algorithms for robust MDPs with non-rectangular uncertainty sets.

Distributionally Robust Deep Q-Learning Policy gradient algorithms for robust MDPs with non-rectangular uncertainty sets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.568230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.568230Z digest=sha256:a80aac1cd4993f95683222c8b9053a75c5fa31fe42602218cda1c8dda39cf124

Observation cc11521d-e1d2-4be7-b7f0-b84df17fb853 · outbound

This paper cites On the efficiency of entropic regularized algorithms for optimal transport.

Distributionally Robust Deep Q-Learning On the efficiency of entropic regularized algorithms for optimal transport

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.547637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:49.722090Z digest=sha256:0752f48f5e83207c5be6068175ba66ee6fec0cce240bdabea12906160259a30d

Observation c7092ac4-6b7b-4725-93a5-afb950f7fa7c · outbound

This paper cites Distri- butionally robust Q-learning.

Distributionally Robust Deep Q-Learning Distri- butionally robust Q-learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:59.293038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:49.840858Z digest=sha256:596244d3718e53dc0f9a9cf85854b0ee55de728ff9cc1acdd702ab16dc23e818

Observation 2e2e4b44-ad3f-4779-b02b-8752d1864697 · outbound

This paper cites Generative model for financial time series trained with MMD using a signature kernel.

Distributionally Robust Deep Q-Learning Generative model for financial time series trained with MMD using a signature kernel

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:49.934758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:49.934758Z digest=sha256:556199d5637a5a6903b92e11d04de77fa1fe26c0e420314793b65f5a991384f5

Observation e7ebd7b1-035f-4e75-9410-b57014f7b327 · outbound

This paper cites Robust MDPs with k-rectangular uncertainty.Mathematics of Operations Research, 41(4):1484–1509, 2016.

Distributionally Robust Deep Q-Learning Robust MDPs with k-rectangular uncertainty.Mathematics of Operations Research, 41(4):1484–1509, 2016

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.911349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.014471Z digest=sha256:e4558c875abbef069b300c7781ecb0f7bfab27a0399c82e98918c7d64ef6aecf

Observation 8e1c0b4b-89f1-4d19-ae29-31fa38fe1ceb · outbound

This paper cites Rusu, Joel Veness, Marc G.

Distributionally Robust Deep Q-Learning Rusu, Joel Veness, Marc G

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.541563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.133260Z digest=sha256:fc870bce9066dd7d53f41551914ef3c71e7f9dc7b128dd6bcd51661006813f86

Observation d33f3249-3bb9-4739-af78-1e972bea358f · outbound

This paper cites Robust SGLD algorithm for solving non-convex distributionally robust optimisation problems.

Distributionally Robust Deep Q-Learning Robust SGLD algorithm for solving non-convex distributionally robust optimisation problems

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.421784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.237029Z digest=sha256:3102007ccc0c7a6443f7d66eb2732785a56f8f15b448988cd1d95035aa35f707

Observation de0a5a18-4943-49b3-87de-d696a478000b · outbound

This paper cites Universal approximation results for neural networks with non-polynomial activation function over non-compact domains.

Distributionally Robust Deep Q-Learning Universal approximation results for neural networks with non-polynomial activation function over non-compact domains

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:50.329080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:50.329080Z digest=sha256:9f65d60291eb31cccf0048dc4bb2922ebf605fc390c2a825ef01e472dc4a47ae

Observation c1a5b965-ff5d-465a-bfc0-3bcb47d4fa5e · outbound

This paper cites Neural networks can detect model-free static arbitrage strategies.

Distributionally Robust Deep Q-Learning Neural networks can detect model-free static arbitrage strategies

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.294955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.440013Z digest=sha256:53542f9f8a43c59b698789d4991bfd3c99f70d3621a79eacbf4ccd647ed89a2c

Observation 8042d1d6-830d-4a55-9209-fbb9db4e4e27 · outbound

This paper cites Robust Q-learning algorithm for markov decision processes under Wasserstein uncertainty.

Distributionally Robust Deep Q-Learning Robust Q-learning algorithm for markov decision processes under Wasserstein uncertainty

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:58.082779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.518778Z digest=sha256:7817de6a6fdce06d1c81a29e062524b2535eced57792796313bd961e8872fddd

Observation a298fc5b-8367-49f9-a359-13bf6d74ecac · outbound

This paper cites Non-concave stochastic optimal control in finite discrete time under model uncertainty.

Distributionally Robust Deep Q-Learning Non-concave stochastic optimal control in finite discrete time under model uncertainty

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:26:53.268456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.619941Z digest=sha256:924fe252944f9947215737cede944bfbef95f85d8e19406f354e9ad87585ff40

Observation d9cd9e9c-8e3a-42e0-a377-cd107430d60b · outbound

This paper cites Markov decision processes under model uncertainty.Mathematical Finance, 33(3):618–665, 2023.

Distributionally Robust Deep Q-Learning Markov decision processes under model uncertainty.Mathematical Finance, 33(3):618–665, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.698534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.701032Z digest=sha256:be1dca0e7ee7075bd9f6d28450129559dee4b224dcfcbc7432d9a3aef8dff125

Observation 2afa2ddc-496e-4315-bb2f-1cf669188761 · outbound

This paper cites Robust control of Markov decision processes with uncertain transition matrices.

Distributionally Robust Deep Q-Learning Robust control of Markov decision processes with uncertain transition matrices

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.465811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.787033Z digest=sha256:d04ec56b4ed59fb6245586b8ca15a8e3571dfea50c071caa17119adfbe8a861c

Observation f2560ea4-7670-4b2b-a2bf-adbff400a0d5 · outbound

This paper cites Introduction to entropic optimal transport.

Distributionally Robust Deep Q-Learning Introduction to entropic optimal transport

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:50.872266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:50.872266Z digest=sha256:d8fec1acde558e17b58cdc5f3007a58eea5656645edd57ad6501fc1b2434a724

Observation b2049884-5728-4488-b3a3-ba753efab79c · outbound

This paper cites Robust reinforcement learning using offline data.

Distributionally Robust Deep Q-Learning Robust reinforcement learning using offline data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.328939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:50.954796Z digest=sha256:662218379434d732c62cf437acf4f4117c5eb33b38b67e89334bdc9d71e0dc21

Observation eea69139-5433-4db7-a707-11bffd83a295 · outbound

This paper cites Approximation theory of the MLP model in neural networks.

Distributionally Robust Deep Q-Learning Approximation theory of the MLP model in neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.165931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.087991Z digest=sha256:e8e3c3f3516a1a486e87f0872ab28dcf9c9499c357217e94257354ce3b196ffd

Observation ddba481d-1872-4c3d-8d0d-08a0d80b2ec0 · outbound

This paper cites Distributionally Robust Optimization: A Review.

Distributionally Robust Deep Q-Learning Distributionally Robust Optimization: A Review

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.197695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.197695Z digest=sha256:9a94553700ad0e21b77c74cd5b70496bbf177b62654fc29ea2810276a6abe0d9

Observation 8cbb9b91-566c-43a2-b474-e6d9ec509828 · outbound

This paper cites Distributionally robust model-based reinforcement learning with large state spaces.

Distributionally Robust Deep Q-Learning Distributionally robust model-based reinforcement learning with large state spaces

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:57.038771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.293275Z digest=sha256:ff5597af038699d1cdda85fce17a0e001c38b85e3ef870ff9c40b18523468b8d

Observation c1040ac8-a747-415d-883d-d2afbf624f0f · outbound

This paper cites Principles of mathematical analysis , volume 3.

Distributionally Robust Deep Q-Learning Principles of mathematical analysis , volume 3

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.883665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.344975Z digest=sha256:15f905ae50901452de0e431a0accdf96cc87b41b59a77b7909677aa76b26c5e5

Observation e4022c2a-03d7-4696-aefe-c7f312d7cf30 · outbound

This paper cites Structural estimation of Markov decision processes.

Distributionally Robust Deep Q-Learning Structural estimation of Markov decision processes

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.753591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.495132Z digest=sha256:af0aaa45479786c08b050f65c474ae913cd375b7d016461f5029fb4f55bd548b

Observation de461997-29b7-4fbc-b4f7-2c936034e59a · outbound

This paper cites Universal approximation using feedforward neural networks: A survey of some existing methods, and some new results.

Distributionally Robust Deep Q-Learning Universal approximation using feedforward neural networks: A survey of some existing methods, and some new results

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.643022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.592726Z digest=sha256:1a5bc9e28c6ddce621506fa74e19085e710712a93d00750f786282381c184b33

Observation e5d9e8fa-928a-4bf0-bce8-48391be90837 · outbound

This paper cites A relationship between arbitrary positive matrices and doubly stochastic matrices.

Distributionally Robust Deep Q-Learning A relationship between arbitrary positive matrices and doubly stochastic matrices

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.488784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.691369Z digest=sha256:480b6fd4c24d8341e4fdb29b2cbc782df855cd49bda33c3134057263d1e6a5b0

Observation 66540415-4f95-4bf3-ab46-6545ddfb849a · outbound

This paper cites Distributionally Robust Reinforcement Learning.

Distributionally Robust Deep Q-Learning Distributionally Robust Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.760331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.760331Z digest=sha256:82bd199c0a246bb9db973c6db2278973d2383409e0bb703a1e1627bba8a8459c

Observation 94800049-d17f-4781-8b6f-e4ea6e729172 · outbound

This paper cites Reinforcement learning: An introduction.

Distributionally Robust Deep Q-Learning Reinforcement learning: An introduction

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.830034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.830034Z digest=sha256:7ad413dd60cb24892ff9b3246e524e7b05580d9544152131098b8d3988514eda

Observation daddacaa-5731-4c90-af39-38eaa04cb3cb · outbound

This paper cites Sinkhorn Divergences for Unbalanced Optimal Transport.

Distributionally Robust Deep Q-Learning Sinkhorn Divergences for Unbalanced Optimal Transport

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:51.926466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:51.926466Z digest=sha256:b8d61926eb9d3c10b090a1c3b1953a0900f2472d8d21d3e28dedff67d0f301c7

Observation 90a94789-bcdc-4323-8563-707e0907b41c · outbound

This paper cites Deep reinforcement learning: From Q-learning to deep Q-learning.

Distributionally Robust Deep Q-Learning Deep reinforcement learning: From Q-learning to deep Q-learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.330805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:51.997657Z digest=sha256:506c1f73f1cbc8cbc1af3049bdf9ede65f0d5ef860de20e3b24c1bf6d6134fd6

Observation 9eb8acb0-7ace-4c7a-a575-38fa53266750 · outbound

This paper cites Deep reinforcement learning with double Q-learning.

Distributionally Robust Deep Q-Learning Deep reinforcement learning with double Q-learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:56.157590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.063423Z digest=sha256:90e136fc8a6ee17cd7877d57c4024783eb8f5f8ac8d709d30c829625db79ae79

Observation c2722181-b1d1-4d24-a021-bab0144a683b · outbound

This paper cites Springer, 2009.

Distributionally Robust Deep Q-Learning Springer, 2009

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.978372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.124396Z digest=sha256:76c6347e4ffdcd09e20288768927463b39d679859dd94cff52b54e9332cb30e0

Observation 8a7301d0-28a0-4060-9ebc-8fe544c6c5b7 · outbound

This paper cites Sinkhorn Distributionally Robust Optimization.

Distributionally Robust Deep Q-Learning Sinkhorn Distributionally Robust Optimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:52.197351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:52.197351Z digest=sha256:54f17b7d59c2bb7c038c2739af492cdac1eccd4770ff56709941a1fee57eb2f0

Observation d1c9673b-498d-493f-aa91-e2d8a0a25d65 · outbound

This paper cites Policy gradient in robust MDPs with global convergence guarantee, 2023.

Distributionally Robust Deep Q-Learning Policy gradient in robust MDPs with global convergence guarantee, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.774906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.283186Z digest=sha256:02193315edeb0fffdabf398a401d92c39b0ee4551c1ea8d15c42e4b983b5aa68

Observation 5a9ea900-7541-46ae-8ad3-2da91d988d08 · outbound

This paper cites A finite sample complexity bound for distribu- tionally robust Q-learning.

Distributionally Robust Deep Q-Learning A finite sample complexity bound for distribu- tionally robust Q-learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.587589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.364238Z digest=sha256:1c9029bbf73bbf8cbddc4091f9ec76f88576e006f4cde3904a34558f1fa77ee4

Observation 1fdd2e0b-7b1d-4ce1-9e14-8bbca4e226fd · outbound

This paper cites Sample complexity of variance-reduced distribu- tionally robust Q-learning.

Distributionally Robust Deep Q-Learning Sample complexity of variance-reduced distribu- tionally robust Q-learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.405979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.436747Z digest=sha256:43bf07dfa56c12e38958b65544f0d0ed3385d23f000a9dc8b45e00e46b95eb17

Observation e46641e0-b583-4a1a-a953-2f322065ca27 · outbound

This paper cites Online robust reinforcement learning with model uncertainty.

Distributionally Robust Deep Q-Learning Online robust reinforcement learning with model uncertainty

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:52.508548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:52.508548Z digest=sha256:8ab937bf5fc12c8d541ff30c387b37869c0410d627645b766b20c0babb1d0e29

Observation a4a41a19-4422-4fd5-ae3a-fdaa96631e48 · outbound

This paper cites Policy gradient method for robust reinforcement learning, 2022.

Distributionally Robust Deep Q-Learning Policy gradient method for robust reinforcement learning, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.215824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.591503Z digest=sha256:78f32bf782bd77fd666ea2638aee2ecf2b696fddd6e4d0e2875e59b1fefcb79b

Observation 0848a997-d20d-44f5-984c-55c6e2a6e9e0 · outbound

This paper cites an unresolved cited work.

Distributionally Robust Deep Q-Learning Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:52.659845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:52.659845Z digest=sha256:de5eb993c02a8164f6247edf2fbf64af92c862205e085b0823eb8fe4319c7401

Observation 33fe3bf7-a23e-473b-bd21-36faa5913b7a · outbound

This paper cites Robust Markov decision processes.

Distributionally Robust Deep Q-Learning Robust Markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:55.057577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.749919Z digest=sha256:3c602c102800233e08c60459fbad2e2a9c1a5d154085b6d41f37e3245758d9f6

Observation 1b3e7aef-70fb-45c3-88bd-798f2076fa29 · outbound

This paper cites Distributionally robust Markov decision processes.

Distributionally Robust Deep Q-Learning Distributionally robust Markov decision processes

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.909557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.836265Z digest=sha256:8d37dae30424942a014b01edb9a704e36aaf48a61f32f858902881b56955af76

Observation fc240353-9188-43fd-8ea5-ac5677b2b03d · outbound

This paper cites A convex optimization approach to distributionally robust Markov decision processes with Wasser- stein distance.

Distributionally Robust Deep Q-Learning A convex optimization approach to distributionally robust Markov decision processes with Wasser- stein distance

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.728296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.904606Z digest=sha256:25a0fc6ee0b7d30ebbc7e4048b67ab37edc86a369407c82a4f517588640a997d

Observation 8158a332-f577-4700-bfb1-88cf205a7ebc · outbound

This paper cites Wasserstein distributionally robust stochastic control: A data-driven approach.

Distributionally Robust Deep Q-Learning Wasserstein distributionally robust stochastic control: A data-driven approach

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.550806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:52.969979Z digest=sha256:f11738f15ceb581c4ea117f99ce613bc78e7c271e2071ed547d135725942f1a0

Observation 1e5f6c2c-bf72-44e7-87f3-7a550d9b32ce · outbound

This paper cites On linear optimization over Wasserstein balls.

Distributionally Robust Deep Q-Learning On linear optimization over Wasserstein balls

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:26:54.369542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:26:53.058645Z digest=sha256:74c6ddf3dd68fcf097be7e0923b9471d0507ede1d61914416c162ad6e3b2c39e

Pith citing papers

Observation 1aaa4507-cffe-4cfc-8c0a-8764e250a890 · inbound

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise cites this paper.

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise Distributionally Robust Deep Q-Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:29:35.432393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:59:25.061726Z digest=sha256:4779324c340627f8cae66d109aa2a58a9a2ba0dadb6d8c6daaee0118260a8a50

Observation 8e06f30b-2c3b-48d9-8ffd-d411166a95ae · inbound

Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making cites this paper.

Robustness in Sequential Decision Making under Evolving Uncertainty: Evidence from High-Frequency Market Making Distributionally Robust Deep Q-Learning

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:57:00.795026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-10T09:56:20.351339Z digest=sha256:01959317135489de048c8049d2f8b7e7b6f0bbe7f0eccdad1931ceb9b7c2d849