Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:00:57.747243Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2412.00985.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T05:00:57.747243Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
95 of 95 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e7163470-73cd-472d-b4a6-f1ebc205c09e · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information End-to-end training of deep visuomotor policies
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbeeea6-e844-492d-99e6-5820f41ad289 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning dex- terous in-hand manipulation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd38f47-686d-40a7-a6b4-632b05f0de02 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9f9c37-b818-4012-99d4-d365c7d89a29 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Deep reinforcement learning for autonomous driving: A survey
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cd51fa3-3e5b-4ea3-b2a2-72417c621d80 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Pomdp-based statistical spoken dialog systems: A review
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90bb442-a7bd-4b7e-8cb8-45907d5e3881 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Informing sequential clinical decision-making through reinforcement learning: an empirical study
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a70d1743-7a2a-4167-b3e5-3c6ae7eb9f07 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information The complexity of markov decision processes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b1aab0-39eb-4063-bc2a-0f7b5c539fe7 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Pac reinforcement learning with rich observations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edf07d0-4530-4dc5-9c96-d945087cfccc · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Sample-e fficient reinforce- ment learning of undercomplete POMDPs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35341c44-0538-44a0-ad0d-25693ee815aa · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A counterexample in stochastic optimum control
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cdf6c5-84e0-402f-add6-a0d965b262bd · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information On the complexity of decentralized decision making and detection problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9392e0-05ed-48ab-a20b-8989b49afb6f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Multi- agent actor-critic for mixed cooperative-competitive environments.Advances in Neural Informa- tion Processing Systems, 30, 2017
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f837fb5-a935-4d6e-b68f-d7ec928e50b6 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6db165-42da-4fcc-b2dc-b38653b2e6c9 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Counterfactual multi-agent policy gradients
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae1370d-cfe2-4e80-9d80-3e49b8411cd9 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Grandmas- ter level in StarCraft II using multi-agent reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209f5775-dff3-4ee1-8fbc-63ef7fd0f64c · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning quadrupedal locomotion over challenging terrain
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b086a5a3-43c3-4730-9c9c-83b87f90fc3f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning robust perceptive locomotion for quadrupedal robots in the wild
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 72e2ad17-e185-48ce-962d-7f900b515742 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning by cheating
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f8827aa1-a12f-4b16-b09e-b5af8efc6b35 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Asymmetric actor critic for image-based robot learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3dbd4a4d-4b10-4558-9917-3494ab24f717 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning in pomdps is sample- efficient with hindsight observability
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4833688b-eb7c-41d0-a7a8-da1d7b233706 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Sample- efficient learning of pomdps with multiple observations in hindsight
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation afb14ff5-9add-461f-830f-dd4432bae525 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning in observable POMDPs, without computationally intractable oracles
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b6818fe-b22c-491d-854c-f95eb52a8bed · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information On oracle-e fficient pac rl with rich observations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 47c406f7-1a2d-4fae-afdd-5786d14ebff9 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Provably e fficient rl with rich observations via latent state decoding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc6b3e90-541e-4674-8d86-9ce17c4e78f7 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 217b5634-c3f1-4990-989d-548706502368 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6220eb8-cd38-4b8b-b43c-fb142e56c383 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Provable reinforce- ment learning with a short-term memory
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd64147e-b6ea-4d4e-9728-8d54b405e90e · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information When is partially observable rein- forcement learning not scary? In Conference on Learning Theory, pages 5175–5220, 2022
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 737e6c1e-7190-412b-909f-d6d980fe8e94 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable multi-agent RL with (quasi-)e fficiency: the blessing of information sharing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c383ef95-b6ce-471b-b251-94451ae6aab7 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Schapire
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48aa0392-b631-4fd2-89ff-f49f5db7b751 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Represent to control partially ob- served systems: Representation learning with provable sample efficiency
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2754bc8-d010-4a96-be31-39ae8a9d93a3 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable RL with b-stability: Unified structural condition and sharp sample-e fficient algorithms
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af9dad91-3fc7-48d2-89f3-f5753195c22f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Reinforcement learning from partial observation: Linear function approximation with provable sample e fficiency
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b66ce82c-c450-40e6-93c4-2050fb6e6a4f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision pro- cesses
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aec8ff2d-a16a-497f-9e24-2e5ac3d3718c · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Optimistic MLE: A generic model-based algorithm for partially observable sequential decision making
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 86df68e4-ed4d-4a4f-8aa1-6c371f8fe8fd · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 72599437-b993-4eb5-a50e-1f38ec65ac56 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Partially observable multi-agent reinforcement learning with information sharing, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 105fcce5-289a-46c5-bcf2-969200142602 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Planning in Observable POMDPs in Quasipolynomial Time
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb4df1d-8d25-4114-9ec0-02605981450a · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Planning and learning in partially observ- able systems via filter stability
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b795815-13fc-490f-a582-00a8188a81bd · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Theoretical hardness and tractability of pomdps in rl with partial online state information, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 311d25fe-b0d2-4841-a9de-df18d7812141 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Leveraging fully observable policies for learning under partial observability
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8343e49c-fd3a-448d-994e-3cd4802dda99 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning to jump from pixels
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dca28fe8-bbb3-4f3e-9302-f672302aeb71 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Tgrl: An algorithm for teacher guided reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 27473edb-659a-4496-8580-a62aae1ad57c · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Asymmetric DQN for partially observable reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73818330-0008-423b-82cd-5428279c758f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Perfectdou: Dominating doudizhu with perfect information distillation.Advances in Neural Information Processing Systems, 35:34954–34965, 2022
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 86c05e3a-ba6c-4130-ad4c-fc2068feb608 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Towards unifying behavioral and response diversity for open-ended learn- ing in zero-sum games
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46cfd868-9448-4408-8a6d-965ecd08c15e · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Unbiased asymmetric reinforcement learning under partial observability
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d18c399a-2656-4fb1-9490-1eb289e362f7 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A deeper understanding of state-based critics in multi-agent reinforcement learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 992bd910-ae36-4806-9b81-d38c65da0108 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information On cen- tralized critics in multi-agent reinforcement learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e1a4770-5d02-4db9-b30e-1e7d478fe43c · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning belief representations for partially observable deep rl
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c202d67a-d224-46c0-8f92-717a8020d41d · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning belief representations for imitation learning in pomdps
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf86327d-f6bd-452b-95fd-7fd78af64a55 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Belief-grounded networks for accelerated robot learning under partial observability
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b8675e2-1378-493b-9b45-0aae522d1133 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892cabc8-9d4a-43ba-b625-39329c5c8d51 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Flow-based recurrent belief state learning for pomdps
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4084d94e-4e65-4da7-8968-b4cec9add4ab · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Belief state actor-critic algorithm from separation principle for POMDP
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2882b484-5216-4df3-82a5-5405accbbe24 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Neural belief states for partially observed domains
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6691a7d0-4711-4ad8-b0b7-68895a1f9f99 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f616ca9-4ad1-4b58-9440-50b8e78a14f2 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Common informa- tion based markov perfect equilibria for stochastic games with asymmetric information: Finite games
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af3b9a8a-b593-4bc1-812a-06b5c13054bd · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Decentralized stochastic con- trol with partial history sharing: A common information approach
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd4d198c-2a00-4d7f-a17b-ade838f12842 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Sample-e fficient reinforcement learning of par- tially observable Markov games
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4df6750a-8ad8-4153-98ff-fc5ce1da62c0 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information When Can We Learn General-Sum Markov Games with a Large Number of Players Sample-Efficiently?
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec2666b-1ae8-49a4-bb75-441e8eb21da6 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A sharp analysis of model-based reinforce- ment learning with self-play
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 38b72e48-4179-4bd1-ad2b-17174936a734 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information V-Learning -- A Simple, Efficient, Decentralized Algorithm for Multiagent RL
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2b1acf5-a996-4570-8c66-3f65afcea8f3 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Algorithmic game theory
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 413df1b0-1e47-4182-8beb-ead88dbc369f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Kakade, and Yishay Mansour
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a866159-e575-4a45-9ed7-802c1031900e · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Common information based markov perfect equilibria for linear-gaussian games with asymmetric information
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cedc9884-654f-48df-a2a5-a4435ef11285 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Poste- rior sampling for competitive rl: Function approximation and partial observation
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fa8a4d42-0c85-45f8-aa69-ebeebc845115 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Learning to communicate with deep multi-agent reinforcement learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2231e7d5-4078-439a-8c17-b7b62cfe3929 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Computationally efficient pac rl in pomdps with latent determinism and conditional embeddings
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e14c806f-8f65-42cf-9715-667326470196 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Actor-critic algorithms
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 82d5fc1e-41a3-4ca6-8134-d17fa4c918e2 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Provably e fficient exploration in policy optimization
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31f7914a-92bf-4466-a3a6-50437b3e9e20 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Optimistic policy optimization with bandit feedback
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c73b445d-dab2-4326-8fb5-54831e4d0d0f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Proximal Policy Optimization Algorithms
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b321497-10f6-4430-a6b4-a709545abb65 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A natural policy gradient
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c907966-e501-4a25-b882-e415f34b9484 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Optimality and approxi- mation with policy gradient methods in Markov decision processes
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 292cda01-7a8f-4ac7-a0f0-cfb0b2c95d91 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Information state embedding in partially observable cooperative multi-agent reinforcement learning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c95e5aa1-8890-4093-9d8c-740e3d16bc7e · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Approximate infor- mation state for approximate planning and reinforcement learning in partially observed sys- tems
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c9bc6d5d-eefc-4818-be95-83c0487b7a49 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Stochastic games with one step delay sharing information pattern with application to power control
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e3bb496-8db5-4244-830e-f8d3f22febbf · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A mea- surement study of internet delay asymmetry
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d2cb1c1-a240-4700-9c95-5033cda07d6b · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Repeated games with incomplete information
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 254b8c83-d1ab-4357-a907-d2f66ae616c0 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Information theory: From coding to learning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 67936a39-66f6-4d20-b77b-b4d56d62b14f · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information On Value Functions and the Agent-Environment Boundary
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c07b385-7de8-4e98-8727-ce3a58b94cd4 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A short note on learning discrete distributions
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2580b214-3fde-481e-aa8e-d04beaee22e3 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A characteri- zation of multiclass learnability, 2022
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cffa59e6-104b-47c8-99c6-0c4b76a227a0 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information A characteriza- tion of multiclass learnability
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88707e5b-c738-43f6-a625-0f666337a471 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Approximately optimal approximate reinforcement learning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation db7590f0-0b5f-43a8-81bb-880ea14462cf · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Convex optimization: Algorithms and complexity
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c550fa58-6a96-4029-9677-f99375c1814a · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Tighter problem-dependent regret bounds in reinforce- ment learning without domain knowledge using value function bounds
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b84e87ff-29a4-42be-843a-e0b026fa63cb · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Reward-free exploration for reinforcement learning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a754894a-2951-40e7-909b-970f2e2adf3b · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Is Q-learning provably efficient? In Advances in Neural Information Processing Systems, pages 4863–4873, 2018
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3b66ef2-81f7-40af-ad94-b46acf6a31e8 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information No-regret learning in convex games
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2822478c-8fe3-4201-995a-0bdb4b02f0e7 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information No-regret learning in bayesian games
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b60dc6e-6f1f-43bb-bd5c-1f1a7e5e29ab · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information Bayes correlated equilibria, no-regret dynamics in Bayesian games, and the price of anarchy
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3edf6af-5c31-413e-a19b-9db9d4b3320c · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information By pluggingLt−1(π) into Equation (C.1), with simple algebric manipulations, we prove that: πt h(·|τh)∝πt−1 h (·|τh)exp ηEsh∼bh(τh) h Qt−1 h (τh,sh,·) i
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1032c0d1-fd32-4015-a79e-d617f6fb4d71 · outbound
Provable Partially Observable Reinforcement Learning with Privileged Information bV (mi⋄πk i )⊙πk −i,G i,h+1 (ch+1) # ,H−h + 1 ) , where the last step is by inductive hypothesis. Now note that for anysh,ph,ah, we have bk−1 h (sh,ah) +Eoh+1∼bJk−1 h (·|sh,ah)
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.