Pith. sign in

Paper Citation Record · LEDGER

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration

As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2506.20307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20307 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:05.049084Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved6
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b64a58-902b-42e7-b8c8-7c39dbb83145 · outbound

This paper cites VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.164194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.751681Z digest=sha256:791209c92b0b0a3cfddce3a292fe1b7fd8ce306c42453382def96c72c31069e1

Observation c2935928-7172-4033-b4de-e0536c1748d6 · outbound

This paper cites Hence the right-hand side of the above inequality is a martingale difference sequence.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hence the right-hand side of the above inequality is a martingale difference sequence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.385296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.038586Z digest=sha256:694ba09504474831e0d7e6924cb3bc3a4af343f37d1d6686724777f94ab8551c

Observation d1430f1a-554f-4981-8768-5fe5da62a566 · outbound

This paper cites Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.784230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.784230Z digest=sha256:485723748f388ab626d8dad90bae55dfe7c9f254ef458a26c43463d4a0948ef0

Observation df7ddbfb-ad7c-403d-8beb-40225c342e4c · outbound

This paper cites Deep reinforcement learning from human preferences.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep reinforcement learning from human preferences

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.103150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.794263Z digest=sha256:cf58801c5c839dc74c082cb74a5ea5be71ec94e1ef03f4143dd024e96d699ef4

Observation 71e633ff-c691-4c30-90cf-0455709637f0 · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Guided cost learning: Deep inverse optimal control via policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.072586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.804541Z digest=sha256:22779272ed2767e186d368e75a9d9bc13a0973ecef51783ba27092d48b521043

Observation 7a4acd9a-846d-455d-91cc-1b8b51483044 · outbound

This paper cites Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.057428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.809354Z digest=sha256:e2451174ea4a00eebbdb8518835f6884dfe6c370cd89f1b04bcc6f5052c308a7

Observation 45e1ea72-f27c-4164-9f2b-1e274f29568d · outbound

This paper cites Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.041045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.813739Z digest=sha256:11b9c9bea065c8392d63513f768bba366b4036337bc6b27c3c911a16046ca45f

Observation cd84756d-95f2-4c82-a4d3-743b27f40256 · outbound

This paper cites Deep Q-learning from Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep Q-learning from Demonstrations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.010041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.827906Z digest=sha256:fec3f9cafaab8611b1e510aefd441a2ad631061acfe6088aba40d187be38d868

Observation b133a3c7-4d75-4bfa-9866-a868c4e96ac3 · outbound

This paper cites Generative adversarial imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.993258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.832292Z digest=sha256:43a68b55c9461b4819d966a0eb09e56f1e2293bb7cd86627e4d075044528d27e

Observation 91b69499-6cc8-42b8-ac7c-3b9f59206ec2 · outbound

This paper cites 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.959186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.841127Z digest=sha256:844b3eb0860e1205c8a54e7718e094fa660902bdd73a8e4ca6e4c77d28173829

Observation c8fc85d8-5ad0-41f3-ad6c-8e10520e6407 · outbound

This paper cites Visual imitation learning with patch rewards.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Visual imitation learning with patch rewards

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.928351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.849975Z digest=sha256:1f9287a9b4ba1d29483a30a4f83b0feebe2d101d9f860e7fa2f08c0b8377211b

Observation 742c03b6-3d2f-41c0-bdf4-0f047d30e529 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Learning: A Modern Introduction Using Convex Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.854645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.854645Z digest=sha256:f08a359205e87c826e19353594fcb283f07f3fb56154bb018e0252d352391b26

Observation 3ff3c01e-d0ab-4d8d-9825-ff9e557cf0af · outbound

This paper cites Online Trajectory Planning in Dynamic Environments for Surgical Task Automation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Trajectory Planning in Dynamic Environments for Surgical Task Automation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.911845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.859792Z digest=sha256:92b562c6241ec7925012bed150cea582f7e98f3f6660a40b96467218a36cc702

Observation da17bb64-0cca-4218-9a21-124f22abb2bf · outbound

This paper cites Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.877453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.868617Z digest=sha256:2922647aea81109869bf5875e1bbf2bed4715f731fbfbdcdd0e1ac839523b1de

Observation f879d5e7-62b6-4b22-ae25-c9e88a38667d · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Alvinn: An autonomous land vehicle in a neural network

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.860689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.873024Z digest=sha256:5d5a6ab8e16c00d7477a50a1d3ee1f7bbb06627787618d7058373130744ee12c

Observation 3c0ea454-4000-404a-bf68-3febb61d8b92 · outbound

This paper cites Eluder dimension and the sample complexity of optimistic ex- ploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Eluder dimension and the sample complexity of optimistic ex- ploration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.829117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.887573Z digest=sha256:ef53d7f3d263e9ff925f995e1f49dbaa52fbe35ca3177d04553f897a4e76b26e

Observation 41ed20c2-c928-4329-931b-d583343495c6 · outbound

This paper cites State entropy maximization with random encoders for efficient exploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy maximization with random encoders for efficient exploration

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.812856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.900895Z digest=sha256:f8cf81549faaeffd2cbad811345c51cf2eb56fafcf41c6f84a6a9cf0234fd90c

Observation 8062e568-92f6-4ff9-93b5-cb188b226393 · outbound

This paper cites Optimistic policy optimization with bandit feedback.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Optimistic policy optimization with bandit feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.797507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.905816Z digest=sha256:d0cf89b1a97e85aa3dc8540ac1cd8f0cc295c52af236334491908e596c3f53c4

Observation 1b72b1f6-3135-4fe2-9106-8d1b85e3984f · outbound

This paper cites Online apprenticeship learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online apprenticeship learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.781367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.910538Z digest=sha256:a1b3ffff3d05891446083829137d1db195358bee5b8c85f5da6384eece2b92c5

Observation 5525c199-3a0c-463a-bfe5-386352dddaef · outbound

This paper cites Error bounds of imitating policies and environments.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Error bounds of imitating policies and environments

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.715637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.930935Z digest=sha256:9ebe2eb46c4556b93163af68b0ef44be331dcf27d6566d5969e3447961b0c4d5

Observation f4a1949e-5356-4d75-8e55-016851457701 · outbound

This paper cites Planning for sample efficient imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Planning for sample efficient imitation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.700048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.935986Z digest=sha256:069b8a10949ab8fd6364cb20d30dffc04b89bbe6ec54672defe97192c0f96c4e

Observation 9301e7f9-9829-4f86-b4f6-69119740d485 · outbound

This paper cites Intrinsic reward driven imitation learning via generative model.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Intrinsic reward driven imitation learning via generative model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.684029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.940955Z digest=sha256:b14820052827bfab246c4ac3ac8bbf537194693a9803bc75f4f81507c1d1cf6a

Observation 41454812-50d6-4cb2-ac40-31fe30f4c44a · outbound

This paper cites Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.668551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.946103Z digest=sha256:af2ad878f6f7ccc18b8ea5993aeba6e533e5afb2235a984d8f0e36b40d6199ca

Observation 2e8b9210-f32b-4cad-8c97-6cdc6ee02110 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.652735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.952083Z digest=sha256:c4dff21fa17aa3ab8ce71bbb6594e066bb689c1c5b515f67d177cd78d3642150

Observation 56e791c1-fa1b-4b82-9e48-a56c187039e8 · outbound

This paper cites Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.636759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.957872Z digest=sha256:2a22371984adad58a1350a74b07ce696cf66f862c5089b52fe488e508a599245

Observation 0a00e1f4-25c7-4b9f-9038-864e591e6e5e · outbound

This paper cites A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:01:05.182636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.962898Z digest=sha256:81d85d7c3381206d8c699f0092dfb20522253fac5ad0bbd9b12e13e40fa2b0da

Observation 460a07f9-20da-4e1a-bd98-1d8dd9cbda12 · outbound

This paper cites Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.620954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.968464Z digest=sha256:ecdd22c93ae3feefcf148856eea541e546b29df126c6d4dc5c96582830d44a22

Observation 29734fda-d086-4df7-8427-4b4a258fd90f · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.604679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.973262Z digest=sha256:611f93c76816a23aa8a3d4d5cdb2695e702e672a782414a94456ca4dbb6350e1

Observation 048c802e-77c7-494e-9aee-26f4ad98270c · outbound

This paper cites Table 3: Comparison of three reward components with three attributes.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 3: Comparison of three reward components with three attributes

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.569475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.983256Z digest=sha256:7fb5ee40035f967e93d9d34575459f7516309b270922ef2d5168d034ad471dec

Observation 6b90e6a3-99bb-4dcd-b744-7602181ce3b6 · outbound

This paper cites 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.551681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.987773Z digest=sha256:4eb643cdfdac6a4504a85f4d9082c99e2068b7a1f794425ae5aab180b009bc70

Observation 24521a19-8289-447d-955c-3a17c41e9cac · outbound

This paper cites VDB constrains the information flow in the discriminator using an information bottleneck.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VDB constrains the information flow in the discriminator using an information bottleneck

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.535824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.992820Z digest=sha256:cd557a2f99dff7229ae010e8500fc8d6955397ba95941fbbca93c879b4004a30

Observation e1069b43-fe6c-4bd7-93d9-ebe3814cbe2b · outbound

This paper cites The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.519977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.997706Z digest=sha256:7c72a9ad50583c0724a683fa8983571efac7dd658ea3488a1db144722584f15b

Observation 8ce758f5-e18f-48de-b356-14e6c64c46bd · outbound

This paper cites State entropy estimate as bonus.Following Seo et al.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy estimate as bonus.Following Seo et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.504115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.002708Z digest=sha256:a95fbe1fee1878da9cf23b5bbb5571ead6efc84c586f358860b93529c0fdb4ff

Observation 626a3192-7096-406e-b367-eccaccf50ef2 · outbound

This paper cites We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.470126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.012779Z digest=sha256:a4a582965fadf3aac83641960ce36365c28cf2a04619ad2130510f805c04acd3

Observation b5b7e079-11c4-4f60-8379-ce67da842a0c · outbound

This paper cites Improve vs Expert.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs Expert

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.453937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.018130Z digest=sha256:f7e13f5e3c2ed5854616bbdf21f3143f72ee2e2007159be1faa4f76eda35e7d8

Observation de71d438-3aa6-43c2-953c-46ad03772d60 · outbound

This paper cites Improve vs GIRIL.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs GIRIL

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.437406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.023291Z digest=sha256:1df210930007ceee0ae972d91d12173e323399cb548960f7c6b5ae074ded439a

Observation 95168769-48c7-4292-8631-a800bebb5d95 · outbound

This paper cites (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.403114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.033529Z digest=sha256:375c79d1e4daffd2a6c693ec8cff060af7607f20e100a0eeae738a5ca1740cf9

Observation b7b8e7e6-2f8c-446e-af65-50d250a2a372 · outbound

This paper cites Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.368525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.043470Z digest=sha256:9e37726d899a1eab9e3671875279390d32a0e382578ce3d4d8051f25d25fd4f2

Observation c65daf9e-c0af-46ef-aa7d-b44bcf0bcbee · outbound

This paper cites Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.352347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.049084Z digest=sha256:7bcc491a977eb55cb04c476dc040ca02138242f82afc32d170808a07ff4e86ad

Observation 15ae85a1-92f2-4ddd-a953-0fecbb8f89a3 · outbound

This paper cites Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.419654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.028464Z digest=sha256:6552bb3309c4ddbac68bb510a80d3e22bd74ef9a8b304ff5cac14b2a7cf5340f

Observation 3dca2e0c-e864-40b9-ab17-07e8e35015c4 · outbound

This paper cites Toward the fundamental limits of imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Toward the fundamental limits of imitation learning

Reference 1988

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.844783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.877390Z digest=sha256:1fa27a504d4888d164dec81b83acf3e941d26b88530279414ed045df33e42883

Observation ed4f5999-4c14-45f7-a129-809e339dc0d1 · outbound

This paper cites Relative entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Relative entropy inverse reinforcement learning

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.149046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.757294Z digest=sha256:b7e67c351cb1b0aae567260932a4a86362d000cabf38b3aa231e10f54544371a

Observation 5d5afa12-317f-4884-ad1f-fee39e977ae9 · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Learning structured output representation using deep conditional generative models

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.765510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.915792Z digest=sha256:88cf7913d7fa69dce641bbc7a87758b24e6a1409c561ca2a7b7c11139e217587

Observation b5aefac0-2d17-4260-a7a9-046d7ec9f654 · outbound

This paper cites and Bagnell, J.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration and Bagnell, J

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.586670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.978061Z digest=sha256:3d376646a86b04a4d67c12bf1af44dbe7ed85b45b7d61ade87d2054a98433b0d

Observation 49dc83bb-418a-48f0-a380-a4f038286589 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Provably efficient reinforcement learning with linear function approximation

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.975912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.836637Z digest=sha256:44cbb071cf2756f5b7678cc8ea329f94a748e1feba58ec5cc69e3c4d9347bf68

Observation d20c8eb2-84bc-45df-a5cb-92e43fa43df2 · outbound

This paper cites OpenAI Gym.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration OpenAI Gym

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.762311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.762311Z digest=sha256:83098a74ca9af429a777a4e5574a41b1b88796de74efccfe91187dfe910a998c

Observation ffb10c27-9f1e-4529-a1b7-1c24fc875d59 · outbound

This paper cites Proximal point imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal point imitation learning

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.731530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.925993Z digest=sha256:26090606d3d38927e3cd04d79b25171e402bde04b18a07ac9651e4154b5bc8e5

Observation 35da6538-f151-47cf-8194-bc640d1ed319 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.893867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.893867Z digest=sha256:ad1e8a098637dfb458af65f9fd353a345386bc6dfa9251753f0c54f342a6eb40

Observation 9b272b49-a52a-4f0a-9472-499bc4ac3db2 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Curiosity-driven exploration by self-supervised prediction

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.895423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.864194Z digest=sha256:ed45e9bdcb094a65b6fbcae864c17bd00793a1eb68f2876a0a3b25f382c97d60

Observation e4ddb88d-b381-4354-b3c4-c767fd83b6c0 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mujoco: A physics engine for model-based control

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.748282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.921365Z digest=sha256:5c1e3f8251bfda196f58ab381c4ad53b6400e1f5a4e20a0dafb86c80b464a4b4

Observation 22552bf1-f381-4fa2-b776-3e639ded67e3 · outbound

This paper cites Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.133533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.768492Z digest=sha256:8d1522f30aa63256663f42e2032c6a13a244f36d65eba1c389081e0d365aae9c

Observation 43d0ada8-eb40-4889-bc2d-e2d7d8a241d2 · outbound

This paper cites End-to- end driving via conditional imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration End-to- end driving via conditional imitation learning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.088173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.800160Z digest=sha256:ffd85632df5fbd12a2f8b06e80259f2e523f153b2713058eedf4bd5c883ad7f8

Observation 8fc4b90e-23d8-42d3-8a74-998af7883bec · outbound

This paper cites On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.300489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.778839Z digest=sha256:27eb07d66a20dbd9dd4a165668fd499ef25c10468ee83be6e5842665ad2531f9

Observation d5d1a19d-f8db-45e8-8aa7-21be4d6cd5cb · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Large-Scale Study of Curiosity-Driven Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.773534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.773534Z digest=sha256:3b0aa6f858aa03e4a5c588fa36280f6ac3586f79a5fc013fe284220ab0b54bfb

Observation e83f2d93-b30d-44a7-afce-4c234c4597cf · outbound

This paper cites Hybrid Inverse Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hybrid Inverse Reinforcement Learning

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.226854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.882451Z digest=sha256:edcda8632a4743a8eeb914287abae7a9571cfb4cdf25415796ffa16ffd669cce

Observation 82a4e234-d90e-4be9-9f52-028324fb64e7 · outbound

This paper cites On Computation and Generalization of Generative Adversarial Imitation Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On Computation and Generalization of Generative Adversarial Imitation Learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.117953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.789475Z digest=sha256:4ba1fed668199710d6b106c89902b95fa8c3b410b09dbddaa4785caee77de657

Observation f1698227-9077-41f4-b4ab-59ce9bea3dab · outbound

This paper cites Imitation Learning from Imperfection: Theoretical Justifications and Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Imitation Learning from Imperfection: Theoretical Justifications and Algorithms

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.943710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.845606Z digest=sha256:54b3857465f064d53606132ce2d3f0cd96aca6ff8414ef4f1b95974f70566555

Observation 47b4d814-94b0-4d36-972e-c16de17b9058 · outbound

This paper cites IQ-Learn: Inverse soft-Q Learning for Imitation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration IQ-Learn: Inverse soft-Q Learning for Imitation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.025396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:04.818299Z digest=sha256:c1a5c2e23e9bcf00ee5c8f76271122e8fa529ccb6e2590252b4e52c5c48b6de9

Observation 8276749f-f8d1-469e-8c3e-e8d28e216c89 · outbound

This paper cites Exploration via Elliptical Episodic Bonuses.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Exploration via Elliptical Episodic Bonuses

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.822802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.822802Z digest=sha256:7db331b60dcb593815f2e591ed0de1639d3796249c0da37174c534743d9ac987

Observation 009e47ab-f081-4445-8d6d-f6bf83f30407 · outbound

This paper cites For a fair comparison, we used an identical policy network for all methods.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration For a fair comparison, we used an identical policy network for all methods

Reference 5184

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.488424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T23:01:05.007817Z digest=sha256:812e07b1381c4fcc77e1178142a432cf6b54523ff6a1d1b804dd60fa100558fd

Pith citing papers

No inbound Pith citation observations are available.