Pith. sign in

Paper Citation Record · LEDGER

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs

As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2411.10906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10906 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:18:52.722594Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1398c259-6c1b-47a2-8ff9-9a44ff679c91 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Near-optimal regret bounds for reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.392987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.540827Z digest=sha256:834a1f6bef81ec9127307cff0e040ee75e6734d73d29bacb2a1e7884d88287e8

Observation a6b2ad8b-b71b-4e4b-9143-12d427be329c · outbound

This paper cites Modular multitask reinforcement learning with policy sketches.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Modular multitask reinforcement learning with policy sketches

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.374662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.546721Z digest=sha256:00e0aad8bb98f9271d83c44f18c2531264f43ee657240cff9e74bd882afddceb

Observation b19f5de2-0bf4-413c-accf-cef2344cd4a5 · outbound

This paper cites Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.357070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.552482Z digest=sha256:e73194c017a629ac29e746b1e6e64914e00db3db159a65997c8a3d39efece235

Observation bf0b6466-2172-4009-8e4e-b2eb003ca2e0 · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.340239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.557621Z digest=sha256:862e55d20efdb243d23fb69b2bb9c2c2f286807c3c78a22af2bcc14f218219bd

Observation e1ca11dc-8f20-4c53-8c7b-fe88ab9a6e1e · outbound

This paper cites Bellemare, Joel Veness, and Michael Bowling.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Bellemare, Joel Veness, and Michael Bowling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.321736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.563350Z digest=sha256:445e563fce13f707d40da9b60f181a9bb195d0d870e2a82e2d81a6f3d976be2c

Observation 97a11bf9-ea1d-41d7-ab13-f446d8389c64 · outbound

This paper cites Vallis, Bruno Lacerda, and Nick Hawes.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Vallis, Bruno Lacerda, and Nick Hawes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.302275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.568988Z digest=sha256:0639d2687768df7376e8ae198b61bffde4e4ba7476e878e86766aba3f8b2fe43

Observation d3c4d07f-28e5-4dea-8ca1-9d678e81d4e3 · outbound

This paper cites Towards deployment-efficient reinforcement learning: Lower bound and optimality.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Towards deployment-efficient reinforcement learning: Lower bound and optimality

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.283165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.574866Z digest=sha256:db74da05bf6e9efb0c9c3fadd621d232b58f722f9218239dd414ce5041c26536

Observation 372f5dd2-d667-4c9a-aed9-75d1857ce6d9 · outbound

This paper cites Model-based reinforcement learning with multinomial logistic function approximation.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Model-based reinforcement learning with multinomial logistic function approximation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.262635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.580669Z digest=sha256:c2044b06be6e1bb5449cc2293bbc004ff3ae4aa37926f65b8e791da4aa8c29b2

Observation f6552027-5548-461a-9955-079db7117c44 · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear markov decision processes.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Nearly minimax optimal reinforcement learning for linear markov decision processes

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.245058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.586857Z digest=sha256:a3aae2607a2c28d09aaffcca04aa698d68f0a15c9e21922fd84da0be4409d135

Observation 9bdbc462-3502-4ced-a7b9-6cc5e73d21a9 · outbound

This paper cites an unresolved cited work.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:18:53.221892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.592143Z digest=sha256:a5d9af4b4b40e1bd8f594b4fd1cf9928807899b1176d0236a39a64db522174d2

Observation f82889a5-f102-4f89-9389-c60329763d65 · outbound

This paper cites Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sample-efficient reinforcement learning is feasible for linearly realizable mdps with limited revisiting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.202779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.597152Z digest=sha256:9729c44d2e34f68fa1a5f8054f11168fab8f371e3b3677b0d34a4888e75e611b

Observation 62607fc6-5394-4e64-bfc2-4c11e4c763b9 · outbound

This paper cites Reinforcement learning and bandits for speech and language processing: Tutorial, review and outlook.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning and bandits for speech and language processing: Tutorial, review and outlook

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.183796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.602688Z digest=sha256:5fbc3ef53250c5c39e36f9b29e03cb4d859711be3ffd75c2555782a13263fbf0

Observation 363e5c45-ff6d-4d6b-84a0-5c2db4407193 · outbound

This paper cites Bandit Algorithms , 2020.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Bandit Algorithms , 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.164415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.608122Z digest=sha256:423148d02e8860a75854da801e0a419de34c4e7706fead6b16eb111fb9e610f1

Observation edf825ab-7dd8-48cd-9259-1bc7e99ec9f3 · outbound

This paper cites Asynchronous Methods for Deep Reinforcement Learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Asynchronous Methods for Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.613740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.613740Z digest=sha256:71a05e622fb5546ab159e4d76a837d0ae08bfff82af2db4b89053c4ddbbbfa84

Observation 3821c233-1aad-46e0-bef1-2349c716c40f · outbound

This paper cites Machado, Marc G.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Machado, Marc G

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.145694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.619774Z digest=sha256:0aae6a0194e3ac080b1748c728ef427e2c4d43a5398c99c81cc513977de44195

Observation 32e0c405-9279-463b-ab75-872609f8c0fc · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Playing Atari with Deep Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.625740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.625740Z digest=sha256:3bfff979bb89935fdaaf7ff8f04bbeafcdd6352ff7ff878d488828ab11e8852f

Observation d3e5d9aa-7d43-4023-9463-b9eaa3702465 · outbound

This paper cites Human-level control through deep reinforcement learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Human-level control through deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.631053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.631053Z digest=sha256:89fbf89375da69cfca93caad2fc68896bce40d6cef4de50a5ad6223b8a817b61

Observation b45d5a89-f0e8-47a8-a0c9-41a8f31d7b08 · outbound

This paper cites Genetic multi-armed bandits: a reinforcement learning approach for discrete optimization via simulation.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Genetic multi-armed bandits: a reinforcement learning approach for discrete optimization via simulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:18:52.825722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.636418Z digest=sha256:785802b933fa7df0a8086af9a3613f867227d8d5ac8045e762e9d934fd5d92a3

Observation 24aae1bb-eed8-4143-8c25-ede472d695c1 · outbound

This paper cites Reinforcement learning in linear mdps: Constant regret and representation selection.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning in linear mdps: Constant regret and representation selection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.112624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.641955Z digest=sha256:5fa206cb4f7750402f07da6b811d798e664e4f9e97052cd427a881e0f28f328a

Observation 83d266a6-519a-4f58-8dbe-b371f0cacf06 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming, 2014.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Markov decision processes: discrete stochastic dynamic programming, 2014

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.093425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.647476Z digest=sha256:66b049dd4203fdfda6aefb199321517bbb24f385581256a584100d66b12e5be8

Observation c93c46c4-b60c-442f-ab8f-d8a4a63c2ca0 · outbound

This paper cites Mastering the game of go without human knowledge.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Mastering the game of go without human knowledge

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.073031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.653917Z digest=sha256:44d72d4a872418c37e26f37027f9300be168ff079bd6bb3f17cb451f0bcb2ea9

Observation c47c58c6-9b23-4118-9462-ebd840b1094e · outbound

This paper cites Sublinear Least-Squares Value Iteration via Locality Sensitive Hashing.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sublinear Least-Squares Value Iteration via Locality Sensitive Hashing

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:18:52.794687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.659457Z digest=sha256:c5cc396c4cc8e6a99bcfef230002dc9ba938e9fdb7434b4667363f5462e43253

Observation d6d8bb08-1128-40a9-a6f7-f0c9672a9860 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.665619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.665619Z digest=sha256:d4a01b87e637b5659ec71062ec87c5448dd502e5814e68d22a65a27bc5968395

Observation bbcda732-6ebe-47c8-b30b-a4ea4e5ee5eb · outbound

This paper cites Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.054990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.670906Z digest=sha256:d98d191ad0c749ff13a255c517e72ec5714232fd0a8ca932db5b63b2f05c40c2

Observation ab10977e-602d-4db6-9050-5247ad61381b · outbound

This paper cites Sketching linear classifiers over data streams.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Sketching linear classifiers over data streams

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.036963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.676065Z digest=sha256:16c509caf656da1426705f5b3570ce375b31c3bc2746cfd79657fce718c79b5a

Observation 2717ec9d-87f1-4481-8d9b-643746ffa28a · outbound

This paper cites How close is the sample covariance matrix to the actual covariance matrix?, 2010.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs How close is the sample covariance matrix to the actual covariance matrix?, 2010

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:53.018862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.681797Z digest=sha256:d55813c39d5bff08783ba7c1db81e61b7e1302b235187e4f81cc68d3a129fec5

Observation b9dfbbbb-9606-49cd-b601-3f8bd5a893f5 · outbound

This paper cites Wagenmaker, Yifang Chen, Max Simchowitz, Simon S.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Wagenmaker, Yifang Chen, Max Simchowitz, Simon S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.999857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.687293Z digest=sha256:3b94095129a4ca96fd32ec31165992ab24e67686da1ae11eb41b1e0ea8726f29

Observation cc067ae4-1ac6-4bc0-890f-d8d15c9095c4 · outbound

This paper cites Woodruff.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Woodruff

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:18:52.693929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:18:52.693929Z digest=sha256:616c75a6c1dccde04f722db4b98b4473ce0b8d3d1c29c1d260b91a083df4bcfc

Observation c4134514-bf53-49fa-abf3-c6881c820739 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation under adaptivity constraints.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Provably efficient reinforcement learning with linear function approximation under adaptivity constraints

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.967458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.699407Z digest=sha256:c7e1f0223cd005167f8a45bea2b199bd98e1a9c526349616bda95bb228ae63ac

Observation 4c7a4f59-bac1-4d04-95b2-77bf3ea2cd6c · outbound

This paper cites A general framework for sequential decision-making under adaptivity constraints.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs A general framework for sequential decision-making under adaptivity constraints

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.949510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.706073Z digest=sha256:353ebb1b86d587d41ccb4473926e6eb73a5b69ae2074ea96272b56f71e549806

Observation dd6ca5f4-986f-4186-9216-2d7eba72a948 · outbound

This paper cites Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.930717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.711923Z digest=sha256:2eeb15a1873df6c1a92068048d12664e478ae62a2cfb4b93bae1c6d895332a5f

Observation 3ec2eaef-735e-4c4c-8eae-2c4b6c9f5d15 · outbound

This paper cites Varshney, and Ashish Jagmohan.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Varshney, and Ashish Jagmohan

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.911774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.717114Z digest=sha256:e0a08b2953787113a4eecf5590c168c8301189ee35bb9a759b8dd676c4bfa20b

Observation 17ccee87-b0e9-4823-b61d-0af3101ee28a · outbound

This paper cites Making linear mdps practical via contrastive representation learning.

Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs Making linear mdps practical via contrastive representation learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:18:52.892212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:18:52.722594Z digest=sha256:001fb1d02a64448142ecf35bd61569b8cc9e21e9025c22f4315f557e765ae581

Pith citing papers

No inbound Pith citation observations are available.