Pith. sign in

Paper Citation Record · LEDGER

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs

As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2411.10906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10906 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:18:52.722594Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1398c259-6c1b-47a2-8ff9-9a44ff679c91 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Near-optimal regret bounds for reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.392987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.540827Z digest=sha256:11a370d983c1d0e601954d906d966df2493be931b463daf6e3ff743343300187

Observation a6b2ad8b-b71b-4e4b-9143-12d427be329c · outbound

This paper cites Modular multitask reinforcement learning with policy sketches.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Modular multitask reinforcement learning with policy sketches

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.374662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.546721Z digest=sha256:f55005780f9077b91ff5aee3ca25ac524c9056c4e31d7fba7de631df583cb56e

Observation b19f5de2-0bf4-413c-accf-cef2344cd4a5 · outbound

This paper cites Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.357070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.552482Z digest=sha256:95cbab3d249a46cef8c1daf80235476853d5280259dac2697bf7df3e38cdf305

Observation bf0b6466-2172-4009-8e4e-b2eb003ca2e0 · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.340239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.557621Z digest=sha256:35e209f47488548cf59d852e6ff38d4bc877a65a3fe2dacccfa2b14b21dab06a

Observation e1ca11dc-8f20-4c53-8c7b-fe88ab9a6e1e · outbound

This paper cites Bellemare, Joel Veness, and Michael Bowling.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Bellemare, Joel Veness, and Michael Bowling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.321736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.563350Z digest=sha256:0cac3858b2ab6a4412b4355aa4bdc6331b422c6fd3955e2c1046be50969ee2c7

Observation 97a11bf9-ea1d-41d7-ab13-f446d8389c64 · outbound

This paper cites Vallis, Bruno Lacerda, and Nick Hawes.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Vallis, Bruno Lacerda, and Nick Hawes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.302275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.568988Z digest=sha256:3216c8f4995b0848f31f70e4c2af2134b5a418a5475da34557e7c0d288460638

Observation d3c4d07f-28e5-4dea-8ca1-9d678e81d4e3 · outbound

This paper cites Towards deployment-efficient reinforcement learning: Lower bound and optimality.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Towards deployment-efficient reinforcement learning: Lower bound and optimality

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.283165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.574866Z digest=sha256:b526e970cacb602f5e09ddeffdda16c7647e0c8520161e5223f30212cba6b48a

Observation 372f5dd2-d667-4c9a-aed9-75d1857ce6d9 · outbound

This paper cites Model-based reinforcement learning with multinomial logistic function approximation.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Model-based reinforcement learning with multinomial logistic function approximation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.262635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.580669Z digest=sha256:289b611d34fde1e989b72cb9f576de559693806a5d996c78e281e543f52f00b1

Observation f6552027-5548-461a-9955-079db7117c44 · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear markov decision processes.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Nearly minimax optimal reinforcement learning for linear markov decision processes

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.245058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.586857Z digest=sha256:a1aa1bd6dd3a9c278eea40dcc6897fd611c3213ef937ca9bd2d4e2aada6e8493

Observation 9bdbc462-3502-4ced-a7b9-6cc5e73d21a9 · outbound

This paper cites an unresolved cited work.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:18:53.221892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.592143Z digest=sha256:b4796745a151ee2ee865bdae9545c41eca427d2a0e7ddd7fdd4a67176c52bf9d

Observation f82889a5-f102-4f89-9389-c60329763d65 · outbound

This paper cites Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.202779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.597152Z digest=sha256:9cd17e5ca3559470b020947afd5435e09d4212ad6a0d7e4479c08cb1b3b2db9d

Observation 62607fc6-5394-4e64-bfc2-4c11e4c763b9 · outbound

This paper cites Reinforcement learning and bandits for speech and language processing: Tutorial, review and outlook.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning and bandits for speech and language processing: Tutorial, review and outlook

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.183796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.602688Z digest=sha256:aadb9c6b8f709db33d9fcd108368dc144b57047c4117c9b4da2859ac74030952

Observation 363e5c45-ff6d-4d6b-84a0-5c2db4407193 · outbound

This paper cites Bandit Algorithms , 2020.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Bandit Algorithms , 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.164415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.608122Z digest=sha256:b00cfa425d532ebf5d1fc875608a4dda8771c3857554bd6594c22b093f35004a

Observation edf825ab-7dd8-48cd-9259-1bc7e99ec9f3 · outbound

This paper cites Asynchronous Methods for Deep Reinforcement Learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Asynchronous Methods for Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.613740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.613740Z digest=sha256:71a05e622fb5546ab159e4d76a837d0ae08bfff82af2db4b89053c4ddbbbfa84

Observation 3821c233-1aad-46e0-bef1-2349c716c40f · outbound

This paper cites Machado, Marc G.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Machado, Marc G

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.145694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.619774Z digest=sha256:5fde0128c046fd8ae2a61f1210c7829d186bd9e31f4b5178b4575b3e4a1f9819

Observation 32e0c405-9279-463b-ab75-872609f8c0fc · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Playing Atari with Deep Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.625740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.625740Z digest=sha256:3bfff979bb89935fdaaf7ff8f04bbeafcdd6352ff7ff878d488828ab11e8852f

Observation d3e5d9aa-7d43-4023-9463-b9eaa3702465 · outbound

This paper cites Human-level control through deep reinforcement learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Human-level control through deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.631053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.631053Z digest=sha256:89fbf89375da69cfca93caad2fc68896bce40d6cef4de50a5ad6223b8a817b61

Observation b45d5a89-f0e8-47a8-a0c9-41a8f31d7b08 · outbound

This paper cites Genetic multi-armed bandits: a reinforcement learning approach for discrete optimization via simulation.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Genetic multi-armed bandits: a reinforcement learning approach for discrete optimization via simulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:18:52.825722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.636418Z digest=sha256:ef680cba74da0a15c9b8c9e326aa1431d47c16178277591b2b8a9b1565224c79

Observation 24aae1bb-eed8-4143-8c25-ede472d695c1 · outbound

This paper cites Reinforcement learning in linear mdps: Constant regret and representation selection.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning in linear mdps: Constant regret and representation selection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.112624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.641955Z digest=sha256:239f08e659e63b2b7c941497cf32442762498d8af3cddf5235bd23e0a0a863b2

Observation 83d266a6-519a-4f58-8dbe-b371f0cacf06 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming, 2014.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Markov decision processes: discrete stochastic dynamic programming, 2014

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.093425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.647476Z digest=sha256:acee6b363cefc0ab46677e339aac5925d719832620d4d146be354f1e1c42b34c

Observation c93c46c4-b60c-442f-ab8f-d8a4a63c2ca0 · outbound

This paper cites Mastering the game of go without human knowledge.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Mastering the game of go without human knowledge

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.073031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.653917Z digest=sha256:839ed64be5c4dd3cd87b8d2a8eb9684929bd3ecccb50000c627e5ea8c7b53444

Observation c47c58c6-9b23-4118-9462-ebd840b1094e · outbound

This paper cites Sublinear Least-Squares Value Iteration via Locality Sensitive Hashing.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sublinear Least-Squares Value Iteration via Locality Sensitive Hashing

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:18:52.794687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.659457Z digest=sha256:1ade18f326fb9979e8365f5eb1a399cef946958b35c7b8a26c27460582885e76

Observation d6d8bb08-1128-40a9-a6f7-f0c9672a9860 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.665619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.665619Z digest=sha256:d4a01b87e637b5659ec71062ec87c5448dd502e5814e68d22a65a27bc5968395

Observation bbcda732-6ebe-47c8-b30b-a4ea4e5ee5eb · outbound

This paper cites Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.054990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.670906Z digest=sha256:f1266302db9b65fc9119b5981596b11806cdb402d75494266d7f0873fcbec67e

Observation ab10977e-602d-4db6-9050-5247ad61381b · outbound

This paper cites Sketching linear classifiers over data streams.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sketching linear classifiers over data streams

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.036963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.676065Z digest=sha256:7a4973290a17e7c49e78d60c2e9ce9809b6f5e451dc2645f1c5a160c44a21d0d

Observation 2717ec9d-87f1-4481-8d9b-643746ffa28a · outbound

This paper cites How close is the sample covariance matrix to the actual covariance matrix?, 2010.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs How close is the sample covariance matrix to the actual covariance matrix?, 2010

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.018862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.681797Z digest=sha256:73d9b224814d8ba0635b140d34fe3da7a48fbc98ae81bbee526ba55655563acb

Observation b9dfbbbb-9606-49cd-b601-3f8bd5a893f5 · outbound

This paper cites Wagenmaker, Yifang Chen, Max Simchowitz, Simon S.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Wagenmaker, Yifang Chen, Max Simchowitz, Simon S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.999857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.687293Z digest=sha256:cd91ddf3e240675043dc60a536af187a4cfc54b96287def95d0d87534808196f

Observation cc067ae4-1ac6-4bc0-890f-d8d15c9095c4 · outbound

This paper cites Woodruff.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Woodruff

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.693929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.693929Z digest=sha256:616c75a6c1dccde04f722db4b98b4473ce0b8d3d1c29c1d260b91a083df4bcfc

Observation c4134514-bf53-49fa-abf3-c6881c820739 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation under adaptivity constraints.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Provably efficient reinforcement learning with linear function approximation under adaptivity constraints

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.967458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.699407Z digest=sha256:78f61730f5adb3151f4d8397defcbf33db816a7d26c0a73304eeb3f6e33a17ab

Observation 4c7a4f59-bac1-4d04-95b2-77bf3ea2cd6c · outbound

This paper cites A general framework for sequential decision-making under adaptivity constraints.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs A general framework for sequential decision-making under adaptivity constraints

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.706073Z digest=sha256:ff0b7d46a95ea88d87db2781182922c0e67222fedf65db41f971f1d63151cb8d

Observation dd6ca5f4-986f-4186-9216-2d7eba72a948 · outbound

This paper cites Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.930717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.711923Z digest=sha256:24af687ecf3db11a5ba754aee1ecee69fe605a27cc39d2c648121457906b638b

Observation 3ec2eaef-735e-4c4c-8eae-2c4b6c9f5d15 · outbound

This paper cites Varshney, and Ashish Jagmohan.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Varshney, and Ashish Jagmohan

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.911774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.717114Z digest=sha256:a1fde0165fb3aa521a74821059883032d710909f6dfbeea0ae91cc70c81841d8

Observation 17ccee87-b0e9-4823-b61d-0af3101ee28a · outbound

This paper cites Making linear mdps practical via contrastive representation learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Making linear mdps practical via contrastive representation learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.892212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.722594Z digest=sha256:bd39ff0424d7632389b81d58229c014755c8fb7cb5c5c0d3b4fb1fde59719522

Pith citing papers

No inbound Pith citation observations are available.