Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:18:52.722594Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2411.10906.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:18:52.722594Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1398c259-6c1b-47a2-8ff9-9a44ff679c91 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Near-optimal regret bounds for reinforcement learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a6b2ad8b-b71b-4e4b-9143-12d427be329c · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Modular multitask reinforcement learning with policy sketches
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b19f5de2-0bf4-413c-accf-cef2344cd4a5 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bf0b6466-2172-4009-8e4e-b2eb003ca2e0 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Logarithmic online regret bounds for undiscounted reinforcement learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e1ca11dc-8f20-4c53-8c7b-fe88ab9a6e1e · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Bellemare, Joel Veness, and Michael Bowling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 97a11bf9-ea1d-41d7-ab13-f446d8389c64 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Vallis, Bruno Lacerda, and Nick Hawes
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d3c4d07f-28e5-4dea-8ca1-9d678e81d4e3 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Towards deployment-efficient reinforcement learning: Lower bound and optimality
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 372f5dd2-d667-4c9a-aed9-75d1857ce6d9 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Model-based reinforcement learning with multinomial logistic function approximation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f6552027-5548-461a-9955-079db7117c44 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Nearly minimax optimal reinforcement learning for linear markov decision processes
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9bdbc462-3502-4ced-a7b9-6cc5e73d21a9 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f82889a5-f102-4f89-9389-c60329763d65 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 62607fc6-5394-4e64-bfc2-4c11e4c763b9 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning and bandits for speech and language processing: Tutorial, review and outlook
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 363e5c45-ff6d-4d6b-84a0-5c2db4407193 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Bandit Algorithms , 2020
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation edf825ab-7dd8-48cd-9259-1bc7e99ec9f3 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Asynchronous Methods for Deep Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3821c233-1aad-46e0-bef1-2349c716c40f · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Machado, Marc G
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 32e0c405-9279-463b-ab75-872609f8c0fc · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Playing Atari with Deep Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e5d9aa-7d43-4023-9463-b9eaa3702465 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Human-level control through deep reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45d5a89-f0e8-47a8-a0c9-41a8f31d7b08 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Genetic multi-armed bandits: a reinforcement learning approach for discrete optimization via simulation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 24aae1bb-eed8-4143-8c25-ede472d695c1 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning in linear mdps: Constant regret and representation selection
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 83d266a6-519a-4f58-8dbe-b371f0cacf06 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Markov decision processes: discrete stochastic dynamic programming, 2014
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c93c46c4-b60c-442f-ab8f-d8a4a63c2ca0 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Mastering the game of go without human knowledge
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c47c58c6-9b23-4118-9462-ebd840b1094e · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sublinear Least-Squares Value Iteration via Locality Sensitive Hashing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d6d8bb08-1128-40a9-a6f7-f0c9672a9860 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcda732-6ebe-47c8-b30b-a4ea4e5ee5eb · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ab10977e-602d-4db6-9050-5247ad61381b · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sketching linear classifiers over data streams
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2717ec9d-87f1-4481-8d9b-643746ffa28a · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs How close is the sample covariance matrix to the actual covariance matrix?, 2010
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b9dfbbbb-9606-49cd-b601-3f8bd5a893f5 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Wagenmaker, Yifang Chen, Max Simchowitz, Simon S
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cc067ae4-1ac6-4bc0-890f-d8d15c9095c4 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Woodruff
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4134514-bf53-49fa-abf3-c6881c820739 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Provably efficient reinforcement learning with linear function approximation under adaptivity constraints
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c7a4f59-bac1-4d04-95b2-77bf3ea2cd6c · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs A general framework for sequential decision-making under adaptivity constraints
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dd6ca5f4-986f-4186-9216-2d7eba72a948 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3ec2eaef-735e-4c4c-8eae-2c4b6c9f5d15 · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Varshney, and Ashish Jagmohan
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 17ccee87-b0e9-4823-b61d-0af3101ee28a · outbound
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Making linear mdps practical via contrastive representation learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.