Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:05.049084Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2506.20307.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:05.049084Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 29b64a58-902b-42e7-b8c8-7c39dbb83145 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c2935928-7172-4033-b4de-e0536c1748d6 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hence the right-hand side of the above inequality is a martingale difference sequence
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d1430f1a-554f-4981-8768-5fe5da62a566 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7ddbfb-ad7c-403d-8beb-40225c342e4c · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep reinforcement learning from human preferences
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71e633ff-c691-4c30-90cf-0455709637f0 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Guided cost learning: Deep inverse optimal control via policy optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7a4acd9a-846d-455d-91cc-1b8b51483044 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45e1ea72-f27c-4164-9f2b-1e274f29568d · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd84756d-95f2-4c82-a4d3-743b27f40256 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep Q-learning from Demonstrations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b133a3c7-4d75-4bfa-9866-a868c4e96ac3 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 91b69499-6cc8-42b8-ac7c-3b9f59206ec2 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8fc85d8-5ad0-41f3-ad6c-8e10520e6407 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Visual imitation learning with patch rewards
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 742c03b6-3d2f-41c0-bdf4-0f047d30e529 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Learning: A Modern Introduction Using Convex Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff3c01e-d0ab-4d8d-9825-ff9e557cf0af · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Trajectory Planning in Dynamic Environments for Surgical Task Automation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation da17bb64-0cca-4218-9a21-124f22abb2bf · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f879d5e7-62b6-4b22-ae25-c9e88a38667d · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Alvinn: An autonomous land vehicle in a neural network
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c0ea454-4000-404a-bf68-3febb61d8b92 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Eluder dimension and the sample complexity of optimistic ex- ploration
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 41ed20c2-c928-4329-931b-d583343495c6 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy maximization with random encoders for efficient exploration
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8062e568-92f6-4ff9-93b5-cb188b226393 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Optimistic policy optimization with bandit feedback
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b72b1f6-3135-4fe2-9106-8d1b85e3984f · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online apprenticeship learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5525c199-3a0c-463a-bfe5-386352dddaef · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Error bounds of imitating policies and environments
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4a1949e-5356-4d75-8e55-016851457701 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Planning for sample efficient imitation learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9301e7f9-9829-4f86-b4f6-69119740d485 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Intrinsic reward driven imitation learning via generative model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 41454812-50d6-4cb2-ac40-31fe30f4c44a · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e8b9210-f32b-4cad-8c97-6cdc6ee02110 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56e791c1-fa1b-4b82-9e48-a56c187039e8 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a00e1f4-25c7-4b9f-9038-864e591e6e5e · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 460a07f9-20da-4e1a-bd98-1d8dd9cbda12 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 29734fda-d086-4df7-8427-4b4a258fd90f · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Maximum entropy inverse reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 048c802e-77c7-494e-9aee-26f4ad98270c · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 3: Comparison of three reward components with three attributes
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b90e6a3-99bb-4dcd-b744-7602181ce3b6 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24521a19-8289-447d-955c-3a17c41e9cac · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VDB constrains the information flow in the discriminator using an information bottleneck
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e1069b43-fe6c-4bd7-93d9-ebe3814cbe2b · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st)
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ce758f5-e18f-48de-b356-14e6c64c46bd · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy estimate as bonus.Following Seo et al
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 626a3192-7096-406e-b367-eccaccf50ef2 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5b7e079-11c4-4f60-8379-ce67da842a0c · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs Expert
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de71d438-3aa6-43c2-953c-46ad03772d60 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs GIRIL
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95168769-48c7-4292-8631-a800bebb5d95 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b7b8e7e6-2f8c-446e-af65-50d250a2a372 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c65daf9e-c0af-46ef-aa7d-b44bcf0bcbee · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...)
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15ae85a1-92f2-4ddd-a953-0fecbb8f89a3 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3dca2e0c-e864-40b9-ab17-07e8e35015c4 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Toward the fundamental limits of imitation learning
Reference 1988
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed4f5999-4c14-45f7-a129-809e339dc0d1 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Relative entropy inverse reinforcement learning
Reference 2002
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d5afa12-317f-4884-ad1f-fee39e977ae9 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Learning structured output representation using deep conditional generative models
Reference 2003
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b5aefac0-2d17-4260-a7a9-046d7ec9f654 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration and Bagnell, J
Reference 2008
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49dc83bb-418a-48f0-a380-a4f038286589 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Provably efficient reinforcement learning with linear function approximation
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d20c8eb2-84bc-45df-a5cb-92e43fa43df2 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration OpenAI Gym
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb10c27-9f1e-4529-a1b7-1c24fc875d59 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal point imitation learning
Reference 2012
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 35da6538-f151-47cf-8194-bc640d1ed319 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal Policy Optimization Algorithms
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b272b49-a52a-4f0a-9472-499bc4ac3db2 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Curiosity-driven exploration by self-supervised prediction
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4ddb88d-b381-4354-b3c4-c767fd83b6c0 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mujoco: A physics engine for model-based control
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22552bf1-f381-4fa2-b776-3e639ded67e3 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 43d0ada8-eb40-4889-bc2d-e2d7d8a241d2 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration End-to- end driving via conditional imitation learning
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8fc4b90e-23d8-42d3-8a74-998af7883bec · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5d1a19d-f8db-45e8-8aa7-21be4d6cd5cb · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Large-Scale Study of Curiosity-Driven Learning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83f2d93-b30d-44a7-afce-4c234c4597cf · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hybrid Inverse Reinforcement Learning
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82a4e234-d90e-4be9-9f52-028324fb64e7 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On Computation and Generalization of Generative Adversarial Imitation Learning
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f1698227-9077-41f4-b4ab-59ce9bea3dab · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Imitation Learning from Imperfection: Theoretical Justifications and Algorithms
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 47b4d814-94b0-4d36-972e-c16de17b9058 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration IQ-Learn: Inverse soft-Q Learning for Imitation
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8276749f-f8d1-469e-8c3e-e8d28e216c89 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Exploration via Elliptical Episodic Bonuses
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009e47ab-f081-4445-8d6d-f6bf83f30407 · outbound
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration For a fair comparison, we used an identical policy network for all methods
Reference 5184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.