Pith. sign in

Paper Citation Record · LEDGER

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:1909.02506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1909.02506 v3

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:54:13.924295Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2da5c4f8-32bb-4c1c-bdd2-7da15c1aafc3 · outbound

This paper cites Schapire.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Schapire

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.513208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.738346Z digest=sha256:b081195d221b81b31db6bda873272ef35b124373c5f60f4740d8033674563494

Observation 2f559d39-bb65-4437-b035-86153a337c30 · outbound

This paper cites Efficient optimal learning for contextual bandits.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Efficient optimal learning for contextual bandits

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.502577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.742927Z digest=sha256:783243c78147c088c9d7c5f8b50f2068906074857a3dbdaf844ad10db49db695

Observation 5765909d-2b49-40bf-9eca-ff4ab1211e87 · outbound

This paper cites Dynamic programming and optimal control , volume 1.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Dynamic programming and optimal control , volume 1

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.491944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.747082Z digest=sha256:79258ca1533d38dffa6e07309477f975b95864fc6a08a456fc79df4c689a43f7

Observation 60f2e34c-8475-4f9c-b361-b942afa71e2e · outbound

This paper cites Approximate Dynamic Programming: Solving the curses of dim ensionality, volume.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Approximate Dynamic Programming: Solving the curses of dim ensionality, volume

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.481694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.751867Z digest=sha256:52dbdd6befc92b19eba67a62364f08fb48e11208d2b4790791c725e96b2b57a4

Observation 9aa0314a-1dfb-46f4-9f31-76c614c0cf63 · outbound

This paper cites Human-level control through deep reinforcement learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Human-level control through deep reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.459429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.760041Z digest=sha256:4a6ef118e013e2236e26d994f133b114ee3aa478617d3dcc8857a385bc3d3f91

Observation a3b5f717-4263-45b2-b6bf-bd38c4a7d0b2 · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Unifying count-based exploration and intrinsic motivation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.448483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.764295Z digest=sha256:28d47937c6b1279530456122c0ea0ed9a73e0f4d2789ffd4e1fb7f39c96949b4

Observation 985da6e6-a9ab-442c-ba7e-e40460a9059c · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Mastering the game of go with deep neural networks and tree search

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.437345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.768222Z digest=sha256:92a6901357f00fbda7a9aba89537a7fedb089c3df828d73afe4f26959a940266

Observation 60df2f6d-336c-4dba-845a-2f26467e27b2 · outbound

This paper cites Ma stering the game of go without human knowledge.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Ma stering the game of go without human knowledge

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.426430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.771912Z digest=sha256:8d92290f718195a0c5df9a1e6190261d3d967c73d88b0200da5e69fe766135b5

Observation ab1668e3-a944-41da-9d4f-f4008e991b56 · outbound

This paper cites Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.415330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.775926Z digest=sha256:a91181c2ad2d0f4ecfd98b503141284685c110b1b1306344e3cce85ab912ae0d

Observation c6d76ad9-7b13-42ce-85ed-7ef61dd5d296 · outbound

This paper cites Is Q-learning provably efficient? In Proceedings of Advances in Neural Information Processing S ystems (NeurIPS) , pages 4863–4873,.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Is Q-learning provably efficient? In Proceedings of Advances in Neural Information Processing S ystems (NeurIPS) , pages 4863–4873,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.404304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.779697Z digest=sha256:67e7f789e7c15ffa07fc3b084c793bfca5c63fe280bf21ac82d5737083d46181

Observation 7541a362-e33a-4c63-b5bc-e4617e947fd7 · outbound

This paper cites Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:54:14.003851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.784020Z digest=sha256:b5496721e7ee66e8f1a2c9b6ca90ec4d2532502ee928861b4a58ba722282d03e

Observation c6ff1699-9e4d-4039-a20c-1bfea857f4f2 · outbound

This paper cites Min imax regret bounds for reinforcement learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Min imax regret bounds for reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.393099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.788388Z digest=sha256:18e001c7e321e486b5c40486ea243059e4faa7fa47cd9b7cd3aafe6ee07cc9d6

Observation 6457e25c-b60e-4e10-aab5-a021349f3c44 · outbound

This paper cites Pac model-free reinforcement learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Pac model-free reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.382243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.792198Z digest=sha256:5bea31373237f9232bf7922d42417b38686699a4d7e6de8b932c6ab9091de565

Observation 53dad104-a1ff-48bf-b438-5895be4eede5 · outbound

This paper cites Speedy q-learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Speedy q-learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.371874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.796245Z digest=sha256:6a6804e0fab688e019d21805cc852c55f8d3d32e2fa217dc0f00c1987c71ab36

Observation 3710c4c3-634f-44ba-8bd7-1d9bd8e7c6f3 · outbound

This paper cites Learning rates for q-lear ning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Learning rates for q-lear ning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.360470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.800280Z digest=sha256:d1ca532a98bd331f59dbe3151aacd28a3c39d9952d9734663e040579080400b5

Observation 012aa6c1-db3b-4f91-b3da-ced392c3c3af · outbound

This paper cites Variance re duced value iteration and faster al- gorithms for solving markov decision processes.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Variance re duced value iteration and faster al- gorithms for solving markov decision processes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.348945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.804599Z digest=sha256:6a56d96d5b28d580044c678589012c0d9e34dfdc2cd25d57799f481419c8a9f5

Observation f99b3821-9f17-4986-b7ab-706eb62a1154 · outbound

This paper cites An analysis of bid-price con trols for network revenue manage- ment.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank An analysis of bid-price con trols for network revenue manage- ment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.337365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.808675Z digest=sha256:b71931bb36fed7e63bd5cae00a21d492fef4b805e6bb8beb31ba57e506bc5a43

Observation 257226e7-a698-4515-b100-85dfb2a514f3 · outbound

This paper cites Dynamic bid prices in revenue management.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Dynamic bid prices in revenue management

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.326612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.812664Z digest=sha256:100329bae3cba130966b44c8528d522a1d8970e5fd5459f1c2a236b45af15e6b

Observation b31dc2dc-f577-4717-b60e-64d439a354da · outbound

This paper cites Asynchronous metho ds for deep reinforcement learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Asynchronous metho ds for deep reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.315679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.816770Z digest=sha256:5bbb3994e953721a53edffca197fc0186eb106c9fc9ee474bbbadde68f77379e

Observation 4a48be75-ec48-4686-883c-ce9ec514f066 · outbound

This paper cites The malmo platform for arti- ficial intelligence experimentation.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank The malmo platform for arti- ficial intelligence experimentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.304884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.820740Z digest=sha256:c75a63ac01f94029d8d63f1ee1f5c9e3e26576d06a50e5e2e36f31e5ed55aa21

Observation a09365b0-df07-499c-9f80-b48ac7c68536 · outbound

This paper cites Gambling in a rigged casino: The adversarial multi-armed bandit problem.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Gambling in a rigged casino: The adversarial multi-armed bandit problem

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.293282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.825045Z digest=sha256:7c20d7c4ae3a5f636cff6eacb71e70126a0e253d27f35ac431c2072e7d512e65

Observation 0aa6861f-a9e4-4934-9c9d-2c2a5a5c2caa · outbound

This paper cites Learning from delayed rewards.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Learning from delayed rewards

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.282303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.829012Z digest=sha256:5f87473a9be231609dd1b9b292c99c4088a5cd9a5d71eba8ccd49f9f710887da

Observation a9b4beea-ee82-4b7a-9ff5-064f5027c09f · outbound

This paper cites Q-learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Q-learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.271279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.832613Z digest=sha256:76a8f45675e4ca7d26308fa5096b9ddb785ad8bd86cc6c6564f1b3be0eaae235

Observation 543ac82f-4627-4459-b740-67a81400fc24 · outbound

This paper cites Asynchronous stochastic approximation and q -learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Asynchronous stochastic approximation and q -learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.259376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.836099Z digest=sha256:7573981769070ee16d0602436c31f1c5f72779bc3fb5e35809fe479ecb8e0f9c

Observation fa802055-7200-4088-b23e-ee942157aadc · outbound

This paper cites Regu- larized policy iteration with nonparametric function spaces.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Regu- larized policy iteration with nonparametric function spaces

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.247160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.839619Z digest=sha256:bb1ceaee820c5a39839c8f19a3a977ae8ab7c6824997a8ac14b1fc3870431192

Observation be1358dd-9d80-4b2a-aa15-aa4a969a9253 · outbound

This paper cites Finite-sample analysis of least- squares policy iteration.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Finite-sample analysis of least- squares policy iteration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.234671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.843708Z digest=sha256:357ce11d7f67e1b764ef493214614853f6f85d2de832243c9587c533010e76f4

Observation 1bdd4641-48d3-470b-b6f8-d777a05c04ac · outbound

This paper cites Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.223709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.847618Z digest=sha256:c4c6f7db6c760f0f9bca14a29a61c8dddf9aa8df4cab0c35e522b5cf703c2594

Observation 3ada3838-90ba-47e3-a4c1-a75e3b3722d7 · outbound

This paper cites Finite-time bounds for fit ted value iteration.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Finite-time bounds for fit ted value iteration

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.212025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.851775Z digest=sha256:d467fab2a1818b2d07f013c8d66c80094383b38a8b97fd6fcb0b00324363e5d0

Observation 056c1862-9c91-4e10-a7de-62d4e0d4e10f · outbound

This paper cites Information-theoretic consideratio ns in batch reinforcement learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Information-theoretic consideratio ns in batch reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.200101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.855880Z digest=sha256:0174009bbc9d6fe451455a3577a23679667cce73ce2d8678a64eb98f0382b427

Observation 0002a454-f72e-466d-ac60-e94f78eff01d · outbound

This paper cites Stable function approximation in dynamic pro gramming.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Stable function approximation in dynamic pro gramming

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.188519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.859818Z digest=sha256:4b9d66fbf6e85b7fb24d1820bddd2171ca9ac95e043f8066a0507f769e8c746a

Observation 70db911d-5bc8-49d7-9b49-3b148c7fdae3 · outbound

This paper cites Provably Efficient Reinforcement Learning with Linear Function Approximation.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Provably Efficient Reinforcement Learning with Linear Function Approximation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:13.863800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:13.863800Z digest=sha256:a3fe7e6c27032ab016b19f4dfa0113072a6bdee3580bb0b67676fa9ae8277a70

Observation 6d6e05e9-b7c9-4ff2-abcb-f222d4938260 · outbound

This paper cites On oracle-efficient pac rl with rich observations.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On oracle-efficient pac rl with rich observations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.176506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.868211Z digest=sha256:a649364f9d22f4ce7d8ba56194f0d7f6fa6171e2aaee6c0b3cbaf61b3809d88d

Observation c6df8ebc-569a-4987-95c3-27c4c4c98a16 · outbound

This paper cites Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, M iroslav Dudik, and John Langford.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, M iroslav Dudik, and John Langford

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.163682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.871945Z digest=sha256:73267772872b3f60a48016b1e21e8d409aff2dbf923d0c2037678197c2915170

Observation 96047cdd-2dd3-4e8d-b827-f1a3f13fc13c · outbound

This paper cites Efficient reinforcement learnin g in deterministic systems with value function generalization.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Efficient reinforcement learnin g in deterministic systems with value function generalization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.151711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.875576Z digest=sha256:a7e5002b1997510813842a01f886255bbb581f196a50dedbc2c1ffd7a9e4ac32

Observation 4ea3b8f1-d38c-4c25-8cb4-a165513d1543 · outbound

This paper cites N ear-optimal time and sample complexities for solving discounted markov decision process with a ge nerative model.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank N ear-optimal time and sample complexities for solving discounted markov decision process with a ge nerative model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.138620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.879925Z digest=sha256:c40388ac435f6e5db5d1ed98c50822759444f33a77989dba07c1f0f5fc4e68cd

Observation 7efce2da-627c-4d07-a012-eb34ddf29243 · outbound

This paper cites Glo bal convergence of policy gradient methods for linearized control problems.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Glo bal convergence of policy gradient methods for linearized control problems

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.125486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.883514Z digest=sha256:637d1030ea1a9ed6c7536b8d871d1c8d32523191917866aac805c45c3e9960b6

Observation 48344bfe-79b9-435f-a0a5-561279b0772f · outbound

This paper cites Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T04:54:13.886912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:54:13.886912Z digest=sha256:20d6f22862712cac460ea6a72ed7ee3520a0ea89c6a6e4b065ab845e289c73ca

Observation b7573e80-b58a-4011-874b-6b5b620b634d · outbound

This paper cites Taming the monster: A fast and simple algorithm for contextual bandits.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Taming the monster: A fast and simple algorithm for contextual bandits

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.112407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.890641Z digest=sha256:bde4c97f7fe0697912803ad9a847e2fed54e6b57132254f12732033080144397

Observation d86242a7-485e-42aa-905c-954229a22d59 · outbound

This paper cites Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T04:54:13.963706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.894260Z digest=sha256:617664a0b681d1ea2a1129e06c0c6cb878919a57b1540c89d0a71464f7cc63c1

Observation 51440e06-2d73-446b-9d59-e3b16f2cbdea · outbound

This paper cites On learning sets and functions.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On learning sets and functions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.101367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.897850Z digest=sha256:e5136c786a19c65270cc8eaeb1bbf24da6a0ecf234de9f79fd1433c064e028ea

Observation d12b1613-a872-4dc4-b587-ae450c1e7b5c · outbound

This paper cites Decision theoretic generalizations of the pac mo del for neural net and other learning applications.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Decision theoretic generalizations of the pac mo del for neural net and other learning applications

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.090701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.901592Z digest=sha256:17bb4fc19451221122a10a8fa60d45da9ee3ba745b7f53e54442085b4f5b9ad4

Observation 05a60069-8a0e-4796-9b03-781f3fdfc24a · outbound

This paper cites Rates of convergence in the central limit the orem for empirical processes.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Rates of convergence in the central limit the orem for empirical processes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.080068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.905224Z digest=sha256:5839b46c18c4d3fe2e05af87de233570c56fe774d17ed437398d4ba823d7037e

Observation c2f41ed6-565d-470f-a3d2-221fc359dfa0 · outbound

This paper cites On general minimax theorems.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On general minimax theorems

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.067730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.909420Z digest=sha256:b4f1d89ff98adff5bc8c94fe04dbeeea0d6cd6c74e398907d67e245614846dd2

Observation 4eb87988-5959-4ab2-b515-148935a61a47 · outbound

This paper cites On minimum volume ellipsoids containing part of a give n ellipsoid.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On minimum volume ellipsoids containing part of a give n ellipsoid

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.054233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.913339Z digest=sha256:c5911f7b5ca0521d4a0e2b73a6f8f34ef9ee3e2aceb9efc2bb7841f81438b93b

Observation c7aa974c-3be4-41e0-9f41-d2f8ae213fef · outbound

This paper cites Sphere packing numbers for subsets of the bo olean n-cube with bounded vapnik- chervonenkis dimension.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Sphere packing numbers for subsets of the bo olean n-cube with bounded vapnik- chervonenkis dimension

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.040507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.916872Z digest=sha256:32b37c709f8b817a3a626475962eb4e1bdad4745a2cb7d1907a69ad0b246f965

Observation 2c84bc67-97d4-4741-84f9-54cdb046dbd1 · outbound

This paper cites Springer Science & Business Media, 2013.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Springer Science & Business Media, 2013

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.027802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.920614Z digest=sha256:8faaa2c4b55ca481ccc59f55c0273b27b232fddd9035fa3c86ae16bb446fceed

Observation aced5f6a-38ae-4bb1-adfe-07eddab1dd2a · outbound

This paper cites Convergence of stochastic processes.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Convergence of stochastic processes

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:54:14.015789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.924295Z digest=sha256:c4caf79381e05e4308b6e0ce5669d13477f2bdcd08dde59942998dfa4c24293c

Observation a1d8be10-b6d3-4b4e-9381-4490a608b6b0 · outbound

This paper cites an unresolved cited work.

$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Unresolved cited work

Reference 703

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:54:14.470456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:54:13.755928Z digest=sha256:917799ef1520a6262785a0877fa1e6a450102b54d597acf8f5884abf3763ea88

Pith citing papers

No inbound Pith citation observations are available.