Pith. sign in

Paper Citation Record · LEDGER

Provable Partially Observable Reinforcement Learning with Privileged Information

As of 12 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2412.00985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00985 v3

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:00:57.747243Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact1
  • verified fuzzy67
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7163470-73cd-472d-b4a6-f1ebc205c09e · outbound

This paper cites End-to-end training of deep visuomotor policies.

Provable Partially Observable Reinforcement Learning with Privileged Information End-to-end training of deep visuomotor policies

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.378425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.378425Z digest=sha256:26889c822e1bb086c07c37f3d83f6418dffe3b0e3fec40a61f17de468f66b258

Observation 5cbeeea6-e844-492d-99e6-5820f41ad289 · outbound

This paper cites Learning dex- terous in-hand manipulation.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning dex- terous in-hand manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.383149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.383149Z digest=sha256:8cf32fd6ab3921805ac6dc68b28a02286b5e080375c789feb06d419460fdab43

Observation 1dd38f47-686d-40a7-a6b4-632b05f0de02 · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Provable Partially Observable Reinforcement Learning with Privileged Information Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.387345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.387345Z digest=sha256:b056a17b3170b77032817b3b54c841ef71305b88b8dfa80dcc7692bc40c80a39

Observation ea9f9c37-b818-4012-99d4-d365c7d89a29 · outbound

This paper cites Deep reinforcement learning for autonomous driving: A survey.

Provable Partially Observable Reinforcement Learning with Privileged Information Deep reinforcement learning for autonomous driving: A survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.392093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.392093Z digest=sha256:733dbe8566de94e480104c17cae6736e7ca735e4e31b4b7336c4780c87fcbac4

Observation 6cd51fa3-3e5b-4ea3-b2a2-72417c621d80 · outbound

This paper cites Pomdp-based statistical spoken dialog systems: A review.

Provable Partially Observable Reinforcement Learning with Privileged Information Pomdp-based statistical spoken dialog systems: A review

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.396226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.396226Z digest=sha256:adc7960aa3a8fb68b21d0a24bd4ddb2489809db3bc22fc5b7e383819ff9329f2

Observation c90bb442-a7bd-4b7e-8cb8-45907d5e3881 · outbound

This paper cites Informing sequential clinical decision-making through reinforcement learning: an empirical study.

Provable Partially Observable Reinforcement Learning with Privileged Information Informing sequential clinical decision-making through reinforcement learning: an empirical study

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.400361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.400361Z digest=sha256:4772d06b05afd65af6af72ae9898a9fbb6eaab895d711be01be8d527ba1e9b4e

Observation a70d1743-7a2a-4167-b3e5-3c6ae7eb9f07 · outbound

This paper cites The complexity of markov decision processes.

Provable Partially Observable Reinforcement Learning with Privileged Information The complexity of markov decision processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.404638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.404638Z digest=sha256:c9a61c468249386badbd0e0f515b5754f15e3e97481df76016ab6c2732e3ea3c

Observation 23b1aab0-39eb-4063-bc2a-0f7b5c539fe7 · outbound

This paper cites Pac reinforcement learning with rich observations.

Provable Partially Observable Reinforcement Learning with Privileged Information Pac reinforcement learning with rich observations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.408509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.408509Z digest=sha256:f19ea82daceb1a4c969dfb5e2b39617c0117fe24e1c7dbefc0dce90fcf1df0ba

Observation 6edf07d0-4530-4dc5-9c96-d945087cfccc · outbound

This paper cites Sample-e fficient reinforce- ment learning of undercomplete POMDPs.

Provable Partially Observable Reinforcement Learning with Privileged Information Sample-e fficient reinforce- ment learning of undercomplete POMDPs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.412288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.412288Z digest=sha256:52fa3161b9498f0205e67daeedcb1e97ff2b6f1d793dddd4d30a846f09e64f24

Observation 35341c44-0538-44a0-ad0d-25693ee815aa · outbound

This paper cites A counterexample in stochastic optimum control.

Provable Partially Observable Reinforcement Learning with Privileged Information A counterexample in stochastic optimum control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.416105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.416105Z digest=sha256:588798c43d5973f5ee285622fad43c8ab6e736e47de4cf2640ec20ff14163f43

Observation 85cdf6c5-84e0-402f-add6-a0d965b262bd · outbound

This paper cites On the complexity of decentralized decision making and detection problems.

Provable Partially Observable Reinforcement Learning with Privileged Information On the complexity of decentralized decision making and detection problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.419988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.419988Z digest=sha256:189edae2f90490625d5610f9f80419d8187c87bd2c0315f9419fbbe91a92d489

Observation 0e9392e0-05ed-48ab-a20b-8989b49afb6f · outbound

This paper cites Multi- agent actor-critic for mixed cooperative-competitive environments.Advances in Neural Informa- tion Processing Systems, 30, 2017.

Provable Partially Observable Reinforcement Learning with Privileged Information Multi- agent actor-critic for mixed cooperative-competitive environments.Advances in Neural Informa- tion Processing Systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.424233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.424233Z digest=sha256:db3ae0bbbebd320e8fe07336c87d60b3ac77e0d7f32a52c6460ae667d0849a85

Observation 4f837fb5-a935-4d6e-b68f-d7ec928e50b6 · outbound

This paper cites QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.428099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.428099Z digest=sha256:6eaa1f7ebe9b13021a9a7cbb5743e58223dcbb142377fb711267049441630fd8

Observation 0b6db165-42da-4fcc-b2dc-b38653b2e6c9 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Provable Partially Observable Reinforcement Learning with Privileged Information Counterfactual multi-agent policy gradients

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.431857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.431857Z digest=sha256:8e9a163c935f812b8fa3f1f2b4fb832406b3b79fcb56c26af4420231259e6596

Observation 4ae1370d-cfe2-4e80-9d80-3e49b8411cd9 · outbound

This paper cites Grandmas- ter level in StarCraft II using multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Grandmas- ter level in StarCraft II using multi-agent reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.435449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.435449Z digest=sha256:e1bda790f9643245a28ba6e3a6fbf1e17f6cbcc09998b9ea5bb5754e07cec628

Observation 209f5775-dff3-4ee1-8fbc-63ef7fd0f64c · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning quadrupedal locomotion over challenging terrain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.439379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.439379Z digest=sha256:6ae9da66ca4859372fe9e2a0bc3c012c4d647ac2d0c2c1730bb00869295d3256

Observation b086a5a3-43c3-4730-9c9c-83b87f90fc3f · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.877864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.443169Z digest=sha256:e6907323a7a08c542a523df503f2f66b51bef949e946e3133712271c8a0b48f1

Observation 72e2ad17-e185-48ce-962d-7f900b515742 · outbound

This paper cites Learning by cheating.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning by cheating

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.865509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.446983Z digest=sha256:3168b8f58e1bfe52b2a5d039f1576a8eb698fc1add5e67198065191da0706b47

Observation f8827aa1-a12f-4b16-b09e-b5af8efc6b35 · outbound

This paper cites Asymmetric actor critic for image-based robot learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Asymmetric actor critic for image-based robot learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.853512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.450732Z digest=sha256:9dd1fc5a3d298da84afd905137df8a334f15e256ff122486eddb189a82c30784

Observation 3dbd4a4d-4b10-4558-9917-3494ab24f717 · outbound

This paper cites Learning in pomdps is sample- efficient with hindsight observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning in pomdps is sample- efficient with hindsight observability

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.841243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.454380Z digest=sha256:90b458e2b642a45ba36bbcb66647dfc4f258271d7019cabaa431603eecf5668f

Observation 4833688b-eb7c-41d0-a7a8-da1d7b233706 · outbound

This paper cites Sample- efficient learning of pomdps with multiple observations in hindsight.

Provable Partially Observable Reinforcement Learning with Privileged Information Sample- efficient learning of pomdps with multiple observations in hindsight

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.828703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.458190Z digest=sha256:264fee8cb146c7e1553c2fc093ce13730fe7666121b21a3a73a7a8738065b7d1

Observation afb14ff5-9add-461f-830f-dd4432bae525 · outbound

This paper cites Learning in observable POMDPs, without computationally intractable oracles.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning in observable POMDPs, without computationally intractable oracles

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.815988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.461960Z digest=sha256:dcba1525e68db56011e03489cef851158e61ceaff096fb00a8e2952de3c819ee

Observation 6b6818fe-b22c-491d-854c-f95eb52a8bed · outbound

This paper cites On oracle-e fficient pac rl with rich observations.

Provable Partially Observable Reinforcement Learning with Privileged Information On oracle-e fficient pac rl with rich observations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.803818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.465780Z digest=sha256:f12dadec69f351d9e988caea5452614576cd3288590b00eb5c4d5a38ba9751dc

Observation 47c406f7-1a2d-4fae-afdd-5786d14ebff9 · outbound

This paper cites Provably e fficient rl with rich observations via latent state decoding.

Provable Partially Observable Reinforcement Learning with Privileged Information Provably e fficient rl with rich observations via latent state decoding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.791388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.469539Z digest=sha256:7184d886bfe13aafb46141e95833027853f59911bc127e8992f691b55464ce79

Observation cc6b3e90-541e-4674-8d86-9ce17c4e78f7 · outbound

This paper cites Kinematic state abstraction and provably efficient rich-observation reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Kinematic state abstraction and provably efficient rich-observation reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.779272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.474307Z digest=sha256:6eea09cb547bb3daf63c43aab18ecc3a8e7d493815dc4d0910cfb4995a4ae5e7

Observation 217b5634-c3f1-4990-989d-548706502368 · outbound

This paper cites Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:00:57.892712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.478097Z digest=sha256:617fd06c8b33969a620980e0aa3b1d0b189c8d2691d967945401df6a41b3dd10

Observation d6220eb8-cd38-4b8b-b43c-fb142e56c383 · outbound

This paper cites Provable reinforce- ment learning with a short-term memory.

Provable Partially Observable Reinforcement Learning with Privileged Information Provable reinforce- ment learning with a short-term memory

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.766737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.482113Z digest=sha256:c980fd7ef07257de1d48c298434f3957b6e88b22dd47eaa6cf31570f2d587107

Observation cd64147e-b6ea-4d4e-9728-8d54b405e90e · outbound

This paper cites When is partially observable rein- forcement learning not scary? In Conference on Learning Theory, pages 5175–5220, 2022.

Provable Partially Observable Reinforcement Learning with Privileged Information When is partially observable rein- forcement learning not scary? In Conference on Learning Theory, pages 5175–5220, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.754554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.486135Z digest=sha256:add5bd94baa75d7ae76142c1febf17abd78325ff96f390e292e132c61db62c59

Observation 737e6c1e-7190-412b-909f-d6d980fe8e94 · outbound

This paper cites Partially observable multi-agent RL with (quasi-)e fficiency: the blessing of information sharing.

Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable multi-agent RL with (quasi-)e fficiency: the blessing of information sharing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.742068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.489987Z digest=sha256:b8649f9cc3de0cd81bd0350bd2de6b4f22e864370f48a2f72773ff3c87d27cb1

Observation c383ef95-b6ce-471b-b251-94451ae6aab7 · outbound

This paper cites Schapire.

Provable Partially Observable Reinforcement Learning with Privileged Information Schapire

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.729516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.493807Z digest=sha256:5f804ce037efe2fe4ac5d96fb9548bcb11c4036a8e66ff9881428fd96b984ec6

Observation 48aa0392-b631-4fd2-89ff-f49f5db7b751 · outbound

This paper cites Represent to control partially ob- served systems: Representation learning with provable sample efficiency.

Provable Partially Observable Reinforcement Learning with Privileged Information Represent to control partially ob- served systems: Representation learning with provable sample efficiency

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.717368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.497654Z digest=sha256:41f6d2193275946177382cdb2b4d415b31327f816fbf03d35260ff943682c6dc

Observation e2754bc8-d010-4a96-be31-39ae8a9d93a3 · outbound

This paper cites Partially observable RL with b-stability: Unified structural condition and sharp sample-e fficient algorithms.

Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable RL with b-stability: Unified structural condition and sharp sample-e fficient algorithms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.705076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.501395Z digest=sha256:4d59b6a281f7f0c762d16c941ce25a5a68497cad259b6d42717f818e33885300

Observation af9dad91-3fc7-48d2-89f3-f5753195c22f · outbound

This paper cites Reinforcement learning from partial observation: Linear function approximation with provable sample e fficiency.

Provable Partially Observable Reinforcement Learning with Privileged Information Reinforcement learning from partial observation: Linear function approximation with provable sample e fficiency

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.692309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.505120Z digest=sha256:222c781fa9aa2209343877c5410e457baf6ae777e142bb2d11265afdd72b0e2f

Observation b66ce82c-c450-40e6-93c4-2050fb6e6a4f · outbound

This paper cites Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision pro- cesses.

Provable Partially Observable Reinforcement Learning with Privileged Information Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision pro- cesses

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.679947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.509825Z digest=sha256:10d679db03557c35d1335295b7b4c4ac170078dc0cc8417c0ee89221e14f09d0

Observation aec8ff2d-a16a-497f-9e24-2e5ac3d3718c · outbound

This paper cites Optimistic MLE: A generic model-based algorithm for partially observable sequential decision making.

Provable Partially Observable Reinforcement Learning with Privileged Information Optimistic MLE: A generic model-based algorithm for partially observable sequential decision making

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.667308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.513652Z digest=sha256:5ccc7e43db60c20dedee9ed1b0e9288c6f6cd529b0c0a56a029d442e154092fb

Observation 86df68e4-ed4d-4a4f-8aa1-6c371f8fe8fd · outbound

This paper cites an unresolved cited work.

Provable Partially Observable Reinforcement Learning with Privileged Information Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T05:00:58.655043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.517747Z digest=sha256:7c6a7769511f87feed58bb13f1bff320ccc04a70e20599b3bd311df773b5a126

Observation 72599437-b993-4eb5-a50e-1f38ec65ac56 · outbound

This paper cites Partially observable multi-agent reinforcement learning with information sharing, 2024.

Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable multi-agent reinforcement learning with information sharing, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.642424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.521399Z digest=sha256:271bcc439b7cb45f612c77aa80e0b596ff33962853a5ae7cb385de957946e9a1

Observation 105fcce5-289a-46c5-bcf2-969200142602 · outbound

This paper cites Planning in Observable POMDPs in Quasipolynomial Time.

Provable Partially Observable Reinforcement Learning with Privileged Information Planning in Observable POMDPs in Quasipolynomial Time

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.525271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.525271Z digest=sha256:501d72da02f28f3797b9e46721b8f2f17aeca27251ebae9a1ba3da582e31bfde

Observation 8cb4df1d-8d25-4114-9ec0-02605981450a · outbound

This paper cites Planning and learning in partially observ- able systems via filter stability.

Provable Partially Observable Reinforcement Learning with Privileged Information Planning and learning in partially observ- able systems via filter stability

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.629812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.529316Z digest=sha256:d97d172428d1098ea1df16d4ad3f2172fd8f7f66bf5f73ab2898eca1a9dc8f82

Observation 4b795815-13fc-490f-a582-00a8188a81bd · outbound

This paper cites Theoretical hardness and tractability of pomdps in rl with partial online state information, 2024.

Provable Partially Observable Reinforcement Learning with Privileged Information Theoretical hardness and tractability of pomdps in rl with partial online state information, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.617712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.533144Z digest=sha256:98d6cc08cc637b1ceb42f9596ee945bae8a3bfde3e30448a90314263b2b378b6

Observation 311d25fe-b0d2-4841-a9de-df18d7812141 · outbound

This paper cites Leveraging fully observable policies for learning under partial observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Leveraging fully observable policies for learning under partial observability

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.605312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.536883Z digest=sha256:ed73f4025f32a35f9dba588524d5a3858d243b2887a14857250863bea19ee562

Observation 8343e49c-fd3a-448d-994e-3cd4802dda99 · outbound

This paper cites Learning to jump from pixels.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning to jump from pixels

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.593020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.540594Z digest=sha256:ea403eb5b0df75a5e1b3942f1d3c7d1a99124139b08c3a3ebf492d646d771fca

Observation dca28fe8-bbb3-4f3e-9302-f672302aeb71 · outbound

This paper cites Tgrl: An algorithm for teacher guided reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Tgrl: An algorithm for teacher guided reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.580773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.544380Z digest=sha256:766e90f51d6c1eb7fbe2f18a6dbd84a313be82c224819d188bbb29767de5175c

Observation 27473edb-659a-4496-8580-a62aae1ad57c · outbound

This paper cites Asymmetric DQN for partially observable reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Asymmetric DQN for partially observable reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.568365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.548314Z digest=sha256:724573af1c841151d2c7fb0d27b89b03ec2449bce93fcf9cd3a624cd53312dbf

Observation 73818330-0008-423b-82cd-5428279c758f · outbound

This paper cites Perfectdou: Dominating doudizhu with perfect information distillation.Advances in Neural Information Processing Systems, 35:34954–34965, 2022.

Provable Partially Observable Reinforcement Learning with Privileged Information Perfectdou: Dominating doudizhu with perfect information distillation.Advances in Neural Information Processing Systems, 35:34954–34965, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.556140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.552052Z digest=sha256:94eff338388fc2949103d8b7e92770b6df23aae609ca0480906b934ee6befc78

Observation 86c05e3a-ba6c-4130-ad4c-fc2068feb608 · outbound

This paper cites Towards unifying behavioral and response diversity for open-ended learn- ing in zero-sum games.

Provable Partially Observable Reinforcement Learning with Privileged Information Towards unifying behavioral and response diversity for open-ended learn- ing in zero-sum games

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.544002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.555842Z digest=sha256:929c57a029c0e171b567704ab82b5cb4f106266dff2a8cdc0df9e3256a2c791d

Observation 46cfd868-9448-4408-8a6d-965ecd08c15e · outbound

This paper cites Unbiased asymmetric reinforcement learning under partial observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Unbiased asymmetric reinforcement learning under partial observability

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.531166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.559621Z digest=sha256:758391ca417df3c30de1705bdf5427d4ead6f8155a5551440e0edfab9c045fe5

Observation d18c399a-2656-4fb1-9490-1eb289e362f7 · outbound

This paper cites A deeper understanding of state-based critics in multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information A deeper understanding of state-based critics in multi-agent reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.518806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.563346Z digest=sha256:7563bfcd44da7a066aeb0fa1447d0fe8623705a6f50490d9ae451196827f1cd7

Observation 992bd910-ae36-4806-9b81-d38c65da0108 · outbound

This paper cites On cen- tralized critics in multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information On cen- tralized critics in multi-agent reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.506483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.566988Z digest=sha256:ace73f025cafc70d28e0abeb2bf9f83448161f25f5f9ded039829b59792a3d55

Observation 2e1a4770-5d02-4db9-b30e-1e7d478fe43c · outbound

This paper cites Learning belief representations for partially observable deep rl.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning belief representations for partially observable deep rl

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.493910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.570730Z digest=sha256:d0e02c5990b7c228c979a437dab89c06245cf49a73a527f743149d5f3e5e70b2

Observation c202d67a-d224-46c0-8f92-717a8020d41d · outbound

This paper cites Learning belief representations for imitation learning in pomdps.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning belief representations for imitation learning in pomdps

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.481653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.575636Z digest=sha256:95d89d656e075ab76924d04dafeb2d0a6ed4cf3103004139e5a941cdc59faea7

Observation bf86327d-f6bd-452b-95fd-7fd78af64a55 · outbound

This paper cites Belief-grounded networks for accelerated robot learning under partial observability.

Provable Partially Observable Reinforcement Learning with Privileged Information Belief-grounded networks for accelerated robot learning under partial observability

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.469046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.579698Z digest=sha256:573ef7b5474cab867307ae3af79256a900c11bc61850e888d5959037c37f0885

Observation 1b8675e2-1378-493b-9b45-0aae522d1133 · outbound

This paper cites Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings.

Provable Partially Observable Reinforcement Learning with Privileged Information Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.583507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.583507Z digest=sha256:52d68325b888741e759b526cd96d47d980b1e35f8acb0900840a9bc84726dd45

Observation 892cabc8-9d4a-43ba-b625-39329c5c8d51 · outbound

This paper cites Flow-based recurrent belief state learning for pomdps.

Provable Partially Observable Reinforcement Learning with Privileged Information Flow-based recurrent belief state learning for pomdps

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.456809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.587635Z digest=sha256:fe7d6f78c2fb701ddbb17adb5d714d4af5d4c59a615528b196a346417c373503

Observation 4084d94e-4e65-4da7-8968-b4cec9add4ab · outbound

This paper cites Belief state actor-critic algorithm from separation principle for POMDP.

Provable Partially Observable Reinforcement Learning with Privileged Information Belief state actor-critic algorithm from separation principle for POMDP

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.444516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.591604Z digest=sha256:ebfd59858fbb5ebc81f78e162981e49a7d6c37d37889da32df71a79e1d982578

Observation 2882b484-5216-4df3-82a5-5405accbbe24 · outbound

This paper cites Neural belief states for partially observed domains.

Provable Partially Observable Reinforcement Learning with Privileged Information Neural belief states for partially observed domains

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.432094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.595393Z digest=sha256:fb612c1b0a621994718feae6a1425c93d68bc7a7a5d2fde5c819f34a25f41cf0

Observation 6691a7d0-4711-4ad8-b0b7-68895a1f9f99 · outbound

This paper cites The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models.

Provable Partially Observable Reinforcement Learning with Privileged Information The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.419475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.599071Z digest=sha256:6dc2b6052ac8ffb599917cbf5b05ab64b2f0756a7ec933dcd70e07d457179dea

Observation 7f616ca9-4ad1-4b58-9440-50b8e78a14f2 · outbound

This paper cites Common informa- tion based markov perfect equilibria for stochastic games with asymmetric information: Finite games.

Provable Partially Observable Reinforcement Learning with Privileged Information Common informa- tion based markov perfect equilibria for stochastic games with asymmetric information: Finite games

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.313480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.602837Z digest=sha256:2ff793adbfc643dc1abbc5863d904f7a19f8093aca01ff3431f9e195d664b79f

Observation af3b9a8a-b593-4bc1-812a-06b5c13054bd · outbound

This paper cites Decentralized stochastic con- trol with partial history sharing: A common information approach.

Provable Partially Observable Reinforcement Learning with Privileged Information Decentralized stochastic con- trol with partial history sharing: A common information approach

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.301288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.606535Z digest=sha256:c5ce81967429e8de570c22984de671a6969171dabb9d98846a93c476ea3bc965

Observation cd4d198c-2a00-4d7f-a17b-ade838f12842 · outbound

This paper cites Sample-e fficient reinforcement learning of par- tially observable Markov games.

Provable Partially Observable Reinforcement Learning with Privileged Information Sample-e fficient reinforcement learning of par- tially observable Markov games

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.289167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.610519Z digest=sha256:36c683102647856fe1490b82e2fe1716e6c52a0f2dc0e9e044ba82cfcb0c936b

Observation 4df6750a-8ad8-4153-98ff-fc5ce1da62c0 · outbound

This paper cites When Can We Learn General-Sum Markov Games with a Large Number of Players Sample-Efficiently?.

Provable Partially Observable Reinforcement Learning with Privileged Information When Can We Learn General-Sum Markov Games with a Large Number of Players Sample-Efficiently?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.614294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.614294Z digest=sha256:06649b25664cb66c225c43483fa581a8e26010fe9501b87e5ffa304f1fd160b2

Observation 6ec2666b-1ae8-49a4-bb75-441e8eb21da6 · outbound

This paper cites A sharp analysis of model-based reinforce- ment learning with self-play.

Provable Partially Observable Reinforcement Learning with Privileged Information A sharp analysis of model-based reinforce- ment learning with self-play

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.275849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.618444Z digest=sha256:7c5b9f3268c5f424c4bb7f4118bdc1e2c682586df46156a2e2cfb553e35639a5

Observation 38b72e48-4179-4bd1-ad2b-17174936a734 · outbound

This paper cites V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL.

Provable Partially Observable Reinforcement Learning with Privileged Information V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.622093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.622093Z digest=sha256:edc6f3be8a19021d4aa714ecc0558381e3ebede98c9d50cf1a889ff91c8874a3

Observation b2b1acf5-a996-4570-8c66-3f65afcea8f3 · outbound

This paper cites Algorithmic game theory.

Provable Partially Observable Reinforcement Learning with Privileged Information Algorithmic game theory

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.263260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.626139Z digest=sha256:7dee4f555d98e93e7aa285eb656a8eb194d164954bbd3fa49fd8634b7a26c3a9

Observation 413df1b0-1e47-4182-8beb-ead88dbc369f · outbound

This paper cites Kakade, and Yishay Mansour.

Provable Partially Observable Reinforcement Learning with Privileged Information Kakade, and Yishay Mansour

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.251251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.629852Z digest=sha256:8f81be5d5f939ab082448a321fa2e2ea5e73b913318aed842c4387f02069d9d0

Observation 4a866159-e575-4a45-9ed7-802c1031900e · outbound

This paper cites Common information based markov perfect equilibria for linear-gaussian games with asymmetric information.

Provable Partially Observable Reinforcement Learning with Privileged Information Common information based markov perfect equilibria for linear-gaussian games with asymmetric information

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.238544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.633736Z digest=sha256:980996d4c7bebb484b7117e01e0cfccae6c18d4713c8c9a3548d4b84c61835ec

Observation cedc9884-654f-48df-a2a5-a4435ef11285 · outbound

This paper cites Poste- rior sampling for competitive rl: Function approximation and partial observation.

Provable Partially Observable Reinforcement Learning with Privileged Information Poste- rior sampling for competitive rl: Function approximation and partial observation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.226186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.637547Z digest=sha256:86299d1dce0e9d32fcf0225c9365f25775b4d3407a562741d852fb3bb2ea0087

Observation fa8a4d42-0c85-45f8-aa69-ebeebc845115 · outbound

This paper cites Learning to communicate with deep multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Learning to communicate with deep multi-agent reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.213731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.641729Z digest=sha256:7ebd2af933fcb95e4a5f77f18161b7c19281ce88a1a502e8f8d566b02a4223e5

Observation 2231e7d5-4078-439a-8c17-b7b62cfe3929 · outbound

This paper cites Computationally efficient pac rl in pomdps with latent determinism and conditional embeddings.

Provable Partially Observable Reinforcement Learning with Privileged Information Computationally efficient pac rl in pomdps with latent determinism and conditional embeddings

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.201238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.645575Z digest=sha256:8a7244471eb715c84ce12d2ab4c16a10392e49884dfcda9c44743a7f1c12fc7f

Observation e14c806f-8f65-42cf-9715-667326470196 · outbound

This paper cites Actor-critic algorithms.

Provable Partially Observable Reinforcement Learning with Privileged Information Actor-critic algorithms

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.186967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.649339Z digest=sha256:8cd244db072813ab72050000273f5ad5bdc562bc5e3793f29661041039cddf99

Observation 82d5fc1e-41a3-4ca6-8134-d17fa4c918e2 · outbound

This paper cites Provably e fficient exploration in policy optimization.

Provable Partially Observable Reinforcement Learning with Privileged Information Provably e fficient exploration in policy optimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.173029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.653116Z digest=sha256:20055848adc85d0ea075c61fa54ab01d82d2fbffc69319d33a9fbb2dc6d0a5f0

Observation 31f7914a-92bf-4466-a3a6-50437b3e9e20 · outbound

This paper cites Optimistic policy optimization with bandit feedback.

Provable Partially Observable Reinforcement Learning with Privileged Information Optimistic policy optimization with bandit feedback

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.159754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.657114Z digest=sha256:feaeaaeed4f21b6efe2308c2b580ec033c90b1510972852914d2da32b5c883f2

Observation c73b445d-dab2-4326-8fb5-54831e4d0d0f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Provable Partially Observable Reinforcement Learning with Privileged Information Proximal Policy Optimization Algorithms

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.660877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.660877Z digest=sha256:718ce0efff90cd9f150d4afd39bd58a056587046bfb58b02931457da7b654fae

Observation 4b321497-10f6-4430-a6b4-a709545abb65 · outbound

This paper cites A natural policy gradient.

Provable Partially Observable Reinforcement Learning with Privileged Information A natural policy gradient

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.146247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.664830Z digest=sha256:b9b85fdccbeb523d79642bd8e160ed3a2577ee18f91b4551bdb9b4a92df066d4

Observation 3c907966-e501-4a25-b882-e415f34b9484 · outbound

This paper cites Optimality and approxi- mation with policy gradient methods in Markov decision processes.

Provable Partially Observable Reinforcement Learning with Privileged Information Optimality and approxi- mation with policy gradient methods in Markov decision processes

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.131943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.668539Z digest=sha256:1ae9c50962e2ba2f106313b8e13c20dccfff8de8602058ed1b2059be11adb7c6

Observation 292cda01-7a8f-4ac7-a0f0-cfb0b2c95d91 · outbound

This paper cites Information state embedding in partially observable cooperative multi-agent reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Information state embedding in partially observable cooperative multi-agent reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.119280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.672466Z digest=sha256:48dcc95fccb761a001d55cad76b2b36f696e8a4ef1cbe74ee5d2552c3f1b53f9

Observation c95e5aa1-8890-4093-9d8c-740e3d16bc7e · outbound

This paper cites Approximate infor- mation state for approximate planning and reinforcement learning in partially observed sys- tems.

Provable Partially Observable Reinforcement Learning with Privileged Information Approximate infor- mation state for approximate planning and reinforcement learning in partially observed sys- tems

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.106613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.676320Z digest=sha256:cfdd7144c2b0ce1f621423a463ca2c7976ee4752e9eb6f82d13fce0f4efbef74

Observation c9bc6d5d-eefc-4818-be95-83c0487b7a49 · outbound

This paper cites Stochastic games with one step delay sharing information pattern with application to power control.

Provable Partially Observable Reinforcement Learning with Privileged Information Stochastic games with one step delay sharing information pattern with application to power control

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.094199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.680053Z digest=sha256:f49dc47077e7377d4108536b5bb61d442895203a1e99c24d3ca06a66e68b31f7

Observation 9e3bb496-8db5-4244-830e-f8d3f22febbf · outbound

This paper cites A mea- surement study of internet delay asymmetry.

Provable Partially Observable Reinforcement Learning with Privileged Information A mea- surement study of internet delay asymmetry

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.081640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.683976Z digest=sha256:3d7b4348c6b44ba1b4f634bd65a7c4a49802706716c5ab3cd2497b114c10faa6

Observation 6d2cb1c1-a240-4700-9c95-5033cda07d6b · outbound

This paper cites Repeated games with incomplete information.

Provable Partially Observable Reinforcement Learning with Privileged Information Repeated games with incomplete information

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.687754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.687754Z digest=sha256:64ba6bda629d6a85f6c96a36b76ad47e4c6dcd70b744705bd0640f338594ab4b

Observation 254b8c83-d1ab-4357-a907-d2f66ae616c0 · outbound

This paper cites Information theory: From coding to learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Information theory: From coding to learning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.060933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.691607Z digest=sha256:165c06c0071fc14abb5462a7ce5c20b7d9259c94319f9385396e581333d9725d

Observation 67936a39-66f6-4d20-b77b-b4d56d62b14f · outbound

This paper cites On Value Functions and the Agent-Environment Boundary.

Provable Partially Observable Reinforcement Learning with Privileged Information On Value Functions and the Agent-Environment Boundary

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.695389Z digest=sha256:361b348218237f50f97aaddb3df479fd4a156f7b263ea1156e31860e8208d5ba

Observation 6c07b385-7de8-4e98-8727-ce3a58b94cd4 · outbound

This paper cites A short note on learning discrete distributions.

Provable Partially Observable Reinforcement Learning with Privileged Information A short note on learning discrete distributions

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.699272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.699272Z digest=sha256:828ef0a00e6d96caefcf94bf2a60c82ea810839157b136511850e4b6653dddc2

Observation 2580b214-3fde-481e-aa8e-d04beaee22e3 · outbound

This paper cites A characteri- zation of multiclass learnability, 2022.

Provable Partially Observable Reinforcement Learning with Privileged Information A characteri- zation of multiclass learnability, 2022

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.048694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.703594Z digest=sha256:fa99ffae85212beba4ca7241a0e75a9fb37896021ae226010b4d7a2a1dce10c4

Observation cffa59e6-104b-47c8-99c6-0c4b76a227a0 · outbound

This paper cites A characteriza- tion of multiclass learnability.

Provable Partially Observable Reinforcement Learning with Privileged Information A characteriza- tion of multiclass learnability

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.036051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.707289Z digest=sha256:b2e36c9aee4f595c561fef5a7c19c1f01d0ecc4b2367c0a74f41e3ccf80f52cb

Observation 88707e5b-c738-43f6-a625-0f666337a471 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Approximately optimal approximate reinforcement learning

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.023529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.711055Z digest=sha256:a5696a2f72c3ed38edb6cbbcdc954ab57432779aec690441a76a44d8733d6d65

Observation db7590f0-0b5f-43a8-81bb-880ea14462cf · outbound

This paper cites Convex optimization: Algorithms and complexity.

Provable Partially Observable Reinforcement Learning with Privileged Information Convex optimization: Algorithms and complexity

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:58.010854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.714868Z digest=sha256:c096d6950d4f1464fdc1390d71c69182f8e998cd46d0eba0252e7db6544a6ece

Observation c550fa58-6a96-4029-9677-f99375c1814a · outbound

This paper cites Tighter problem-dependent regret bounds in reinforce- ment learning without domain knowledge using value function bounds.

Provable Partially Observable Reinforcement Learning with Privileged Information Tighter problem-dependent regret bounds in reinforce- ment learning without domain knowledge using value function bounds

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.997919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.718762Z digest=sha256:81c2f9db95cef2af5e661bd207ba88c2a04156468a50f090c03294e796586d15

Observation b84e87ff-29a4-42be-843a-e0b026fa63cb · outbound

This paper cites Reward-free exploration for reinforcement learning.

Provable Partially Observable Reinforcement Learning with Privileged Information Reward-free exploration for reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.984880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.722481Z digest=sha256:d88e414bc725c8a393c05dccaaf190ce52761751b62d01553e2872ee17bb4a6f

Observation a754894a-2951-40e7-909b-970f2e2adf3b · outbound

This paper cites Is Q-learning provably efficient? In Advances in Neural Information Processing Systems, pages 4863–4873, 2018.

Provable Partially Observable Reinforcement Learning with Privileged Information Is Q-learning provably efficient? In Advances in Neural Information Processing Systems, pages 4863–4873, 2018

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.971151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.726414Z digest=sha256:04047a2be08200b3167abf66abe4e6d28849b029ec375d8492383d31a16351c4

Observation c3b66ef2-81f7-40af-ad94-b46acf6a31e8 · outbound

This paper cites No-regret learning in convex games.

Provable Partially Observable Reinforcement Learning with Privileged Information No-regret learning in convex games

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.958482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.730158Z digest=sha256:f067079883ccea2498234dadfdaf6e5ddcb88f88396e1d5a7b504906829ec3a4

Observation 2822478c-8fe3-4201-995a-0bdb4b02f0e7 · outbound

This paper cites No-regret learning in bayesian games.

Provable Partially Observable Reinforcement Learning with Privileged Information No-regret learning in bayesian games

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.945880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.734176Z digest=sha256:b35e250818168671abd6d222718db71d043508dc18ffc1b91c726e93d93456f3

Observation 1b60dc6e-6f1f-43bb-bd5c-1f1a7e5e29ab · outbound

This paper cites Bayes correlated equilibria, no-regret dynamics in Bayesian games, and the price of anarchy.

Provable Partially Observable Reinforcement Learning with Privileged Information Bayes correlated equilibria, no-regret dynamics in Bayesian games, and the price of anarchy

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T05:00:57.737884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:00:57.737884Z digest=sha256:c4ad69961707ef9428610ca3f6c141cf593377f63ccdb4a66aadd5564316325e

Observation c3edf6af-5c31-413e-a19b-9db9d4b3320c · outbound

This paper cites By pluggingLt−1(π) into Equation (C.1), with simple algebric manipulations, we prove that: πt h(·|τh)∝πt−1 h (·|τh)exp ηEsh∼bh(τh) h Qt−1 h (τh,sh,·) i.

Provable Partially Observable Reinforcement Learning with Privileged Information By pluggingLt−1(π) into Equation (C.1), with simple algebric manipulations, we prove that: πt h(·|τh)∝πt−1 h (·|τh)exp ηEsh∼bh(τh) h Qt−1 h (τh,sh,·) i

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:00:57.933014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.742983Z digest=sha256:e2ccf45f4b0bd06c96f43bc62a0fe10d9bd34cb683d0c92e3ffe5b77bca288a8

Observation 1032c0d1-fd32-4015-a79e-d617f6fb4d71 · outbound

This paper cites bV (mi⋄πk i )⊙πk −i,G i,h+1 (ch+1) # ,H−h + 1 ) , where the last step is by inductive hypothesis. Now note that for anysh,ph,ah, we have bk−1 h (sh,ah) +Eoh+1∼bJk−1 h (·|sh,ah).

Provable Partially Observable Reinforcement Learning with Privileged Information bV (mi⋄πk i )⊙πk −i,G i,h+1 (ch+1) # ,H−h + 1 ) , where the last step is by inductive hypothesis. Now note that for anysh,ph,ah, we have bk−1 h (sh,ah) +Eoh+1∼bJk−1 h (·|sh,ah)

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T05:00:57.920214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T05:00:57.747243Z digest=sha256:51125d9906aba0ce3a70359fec7c9ff56a442d77ef3365155c4476182f1d2685

Pith citing papers

No inbound Pith citation observations are available.