Pith. sign in

Paper Citation Record · LEDGER

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration

As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2506.20307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20307 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:05.049084Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved6
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b64a58-902b-42e7-b8c8-7c39dbb83145 · outbound

This paper cites VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VO Q L: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.164194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.751681Z digest=sha256:2cb2e1d9a0607362bd5937e5b0c6033241ceff73f9c7934ae981974bb612c504

Observation c2935928-7172-4033-b4de-e0536c1748d6 · outbound

This paper cites Hence the right-hand side of the above inequality is a martingale difference sequence.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hence the right-hand side of the above inequality is a martingale difference sequence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.385296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.038586Z digest=sha256:6e5200eac649b7dfb6ed38555edac664e92d0265cf8a06ce2547a0af5f37f041

Observation d1430f1a-554f-4981-8768-5fe5da62a566 · outbound

This paper cites Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.784230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.784230Z digest=sha256:cf7f06f0e574a109b1d827b3a8451e97cb217a4bf3f1d296a99fab0e85db236a

Observation df7ddbfb-ad7c-403d-8beb-40225c342e4c · outbound

This paper cites Deep reinforcement learning from human preferences.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep reinforcement learning from human preferences

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.103150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.794263Z digest=sha256:fb55a5792a5a781b168605607eabd3cc71c74d88a8e97ef7ae3182d43911def5

Observation 71e633ff-c691-4c30-90cf-0455709637f0 · outbound

This paper cites Guided cost learning: Deep inverse optimal control via policy optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Guided cost learning: Deep inverse optimal control via policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.072586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.804541Z digest=sha256:7d7375be8bd9165605be2c543e4c5f2c45ef9762da7203ca031cc33c794321f2

Observation 7a4acd9a-846d-455d-91cc-1b8b51483044 · outbound

This paper cites Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Efficient Bias-Span-Constrained Exploration-Exploitation in Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.057428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.809354Z digest=sha256:b185e8d6267b140da0a8b2bc8038452dc370d42ef2aef5e2e552fd257ceee27c

Observation 45e1ea72-f27c-4164-9f2b-1e274f29568d · outbound

This paper cites Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Ess- InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.041045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.813739Z digest=sha256:39c42a75153dae6022e9f23aac94ee07b59a4f5f88f95368f169683f60391559

Observation cd84756d-95f2-4c82-a4d3-743b27f40256 · outbound

This paper cites Deep Q-learning from Demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep Q-learning from Demonstrations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.010041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.827906Z digest=sha256:2d64a605633c2219ce45f9eb04c9c88a0642bd30cadb7f9cb9e69850165e217d

Observation b133a3c7-4d75-4bfa-9866-a868c4e96ac3 · outbound

This paper cites Generative adversarial imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.993258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.832292Z digest=sha256:2d2aea339635557970fc2cdf63266fe7b48e0272dc27f1c4bf33b4b066d50ac5

Observation 91b69499-6cc8-42b8-ac7c-3b9f59206ec2 · outbound

This paper cites 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.959186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.841127Z digest=sha256:9c80d197431807a6c552bcd38c94a38657cfd306e27167c257cf698b0eab99d6

Observation c8fc85d8-5ad0-41f3-ad6c-8e10520e6407 · outbound

This paper cites Visual imitation learning with patch rewards.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Visual imitation learning with patch rewards

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.928351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.849975Z digest=sha256:edb1f6cc6c1351fe299450d5b2f4e0c9c18910172756d621e6804b33b16abe02

Observation 742c03b6-3d2f-41c0-bdf4-0f047d30e529 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Learning: A Modern Introduction Using Convex Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.854645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.854645Z digest=sha256:2fa8e2b0af63c03abf1bb911d9c7f4e7176fc54fc3116fdbab654a049aca5bcc

Observation 3ff3c01e-d0ab-4d8d-9825-ff9e557cf0af · outbound

This paper cites Online Trajectory Planning in Dynamic Environments for Surgical Task Automation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online Trajectory Planning in Dynamic Environments for Surgical Task Automation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.911845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.859792Z digest=sha256:8b0de62ef83f89e75b480c68ff82f3a255ccdd9e71aac4cefc60bfc13c398173

Observation da17bb64-0cca-4218-9a21-124f22abb2bf · outbound

This paper cites Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Variational discriminator bottleneck: Improving imitation learning, inverse RL, and GANs by constraining information flow

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.877453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.868617Z digest=sha256:0ca19b44f0e51de8c6a08f1ce0e4e60cf3a7fa05c4b58cc2849886fdd25b7307

Observation f879d5e7-62b6-4b22-ae25-c9e88a38667d · outbound

This paper cites Alvinn: An autonomous land vehicle in a neural network.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Alvinn: An autonomous land vehicle in a neural network

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.860689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.873024Z digest=sha256:26be7470e8e743bc477f4dc5f89be50091cc0d9fe8a90022a074ccd2771b1a35

Observation 3c0ea454-4000-404a-bf68-3febb61d8b92 · outbound

This paper cites Eluder dimension and the sample complexity of optimistic ex- ploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Eluder dimension and the sample complexity of optimistic ex- ploration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.829117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.887573Z digest=sha256:bc73f061ea49a674cd643b7e6b3f903d91fa3bc78e0628c27590f60c29867fc5

Observation 41ed20c2-c928-4329-931b-d583343495c6 · outbound

This paper cites State entropy maximization with random encoders for efficient exploration.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy maximization with random encoders for efficient exploration

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.812856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.900895Z digest=sha256:194d9ff36f0506bd0c45a1b673c20158691312d117b0b58fd886988ea1d01a8e

Observation 8062e568-92f6-4ff9-93b5-cb188b226393 · outbound

This paper cites Optimistic policy optimization with bandit feedback.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Optimistic policy optimization with bandit feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.797507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.905816Z digest=sha256:55f406b5cf338208d39f3fc64509bef7b029c04a69c05668d9d857fdaa706ebc

Observation 1b72b1f6-3135-4fe2-9106-8d1b85e3984f · outbound

This paper cites Online apprenticeship learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Online apprenticeship learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.781367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.910538Z digest=sha256:dc89a7c59f0f1a0815721ea0650ab0907100dc67ab21a91cf9e8efa38e8a78b4

Observation 5525c199-3a0c-463a-bfe5-386352dddaef · outbound

This paper cites Error bounds of imitating policies and environments.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Error bounds of imitating policies and environments

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.715637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.930935Z digest=sha256:30e0c40c8923df1988e230896d926621ea9a88c718a863c47cc0bea5b18d9548

Observation f4a1949e-5356-4d75-8e55-016851457701 · outbound

This paper cites Planning for sample efficient imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Planning for sample efficient imitation learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.700048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.935986Z digest=sha256:029d0427bb6af1290c5d382bead53bb2d690a0221b239a4e2f161ad4680dbf6d

Observation 9301e7f9-9829-4f86-b4f6-69119740d485 · outbound

This paper cites Intrinsic reward driven imitation learning via generative model.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Intrinsic reward driven imitation learning via generative model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.684029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.940955Z digest=sha256:13e7cf5e003841c0bae929c6154ee0362cef0efdfbc4f9278d66df62c3aa1134

Observation 41454812-50d6-4cb2-ac40-31fe30f4c44a · outbound

This paper cites Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.668551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.946103Z digest=sha256:c07c1d03a275859123161947145589f9927638d54843018ad50c60b320de7019

Observation 2e8b9210-f32b-4cad-8c97-6cdc6ee02110 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.652735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.952083Z digest=sha256:26803aa15792d6dca315051c1a5555a7cf5df65201882d703cda6f6fdfcc7802

Observation 56e791c1-fa1b-4b82-9e48-a56c187039e8 · outbound

This paper cites Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.636759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.957872Z digest=sha256:ad7b3d6dcabc17d16dee80441470256337b0734d83003abbebfcc2039fcd002a

Observation 0a00e1f4-25c7-4b9f-9038-864e591e6e5e · outbound

This paper cites A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation.arXiv preprint arXiv:2311.15238,

Reference 43

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:01:05.182636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.962898Z digest=sha256:51c62408beb31a7330cea85c65fa4cd7afc2c91860ae19a129f1fa2b922adccf

Observation 460a07f9-20da-4e1a-bd98-1d8dd9cbda12 · outbound

This paper cites Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Self-adaptive imitation learning: Learning tasks with delayed rewards from sub-optimal demonstrations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.620954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.968464Z digest=sha256:dc20a38e5129c3cda9ee68e9f550573a9339805119cf61602e8f98f92dac7097

Observation 29734fda-d086-4df7-8427-4b4a258fd90f · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.604679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.973262Z digest=sha256:d7c95b2f3bfa6b7cecf0a2304c0bf167bf5425624181fecc15c2a32ea7670fd1

Observation 048c802e-77c7-494e-9aee-26f4ad98270c · outbound

This paper cites Table 3: Comparison of three reward components with three attributes.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 3: Comparison of three reward components with three attributes

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.569475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.983256Z digest=sha256:3a424a7e07d59dd3071e25c31e1a94768b1b163a47ef490f89cabf4489e60830

Observation 6b90e6a3-99bb-4dcd-b744-7602181ce3b6 · outbound

This paper cites 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration 14 Published as a conference paper at ICLR 2025 Table 4: Demonstration lengths in the Atari environment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.551681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.987773Z digest=sha256:e37aed553949345664466c2ced05154d2f3151377e1d52f5c50b6f0a6b951cb6

Observation 24521a19-8289-447d-955c-3a17c41e9cac · outbound

This paper cites VDB constrains the information flow in the discriminator using an information bottleneck.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration VDB constrains the information flow in the discriminator using an information bottleneck

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.535824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.992820Z digest=sha256:5520c0a2e18526842ca6ce271179550ba247732781d6a2370c26bd70f75e159d

Observation e1069b43-fe6c-4bd7-93d9-ebe3814cbe2b · outbound

This paper cites The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration The module is composed of several neural networks, including recognition network qϕ(z|st, st+1), a generative network pθ(st+1|z, st), and prior network pθ(z|st)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.519977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.997706Z digest=sha256:e411c2c7e52a011ac443e2d7916755205d840a8bde94dfd87c19eb5cf3da8b1c

Observation 8ce758f5-e18f-48de-b356-14e6c64c46bd · outbound

This paper cites State entropy estimate as bonus.Following Seo et al.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration State entropy estimate as bonus.Following Seo et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.504115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.002708Z digest=sha256:d292e77e83a72bfa171e9aa4704d71c64dc93d886119e8bd92d56f62fe62ae69

Observation 626a3192-7096-406e-b367-eccaccf50ef2 · outbound

This paper cites We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration We use the same hyperparameter setting for the different experiments within a domain (Atari vs MuJoCo), apart from a multiplicative constant based on the range of the curiosity

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.470126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.012779Z digest=sha256:a71edd43ff901b71ab1484eae952f4e4f42ad8f323d722237218e58ece8bf5af

Observation b5b7e079-11c4-4f60-8379-ce67da842a0c · outbound

This paper cites Improve vs Expert.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs Expert

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.453937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.018130Z digest=sha256:77fa9a9d549e6a7d2b63c61979e9468966130a80734055ae5d5e33683f1c4c6f

Observation de71d438-3aa6-43c2-953c-46ad03772d60 · outbound

This paper cites Improve vs GIRIL.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Improve vs GIRIL

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.437406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.023291Z digest=sha256:af17accae658ef89af7c55897ceb7a44ff4ae1b7332a81efa9f1841fd47c9947

Observation 95168769-48c7-4292-8631-a800bebb5d95 · outbound

This paper cites (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration (2023), and (3) the analysis leads to a sublinear batch-regret, which is stronger than a sample complexity bound

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.403114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.033529Z digest=sha256:496e8d3e26f40847d2a15c4ed275cd69c5f9cba05ecbff121e6245a9bd4b2253

Observation b7b8e7e6-2f8c-446e-af65-50d250a2a372 · outbound

This paper cites Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Then with probability at least1−2δ, KX k=1 J(π ∗,erk)−J(π,er k) ≤Hlog|A|/η+ηH 3K/2 + 2HK· p 8H 2 log(H· NF (ϵF )/δ) + 4ϵF N+γ·O s H N · H 2 +γ γ dimN (Fh) log(1/δ) !

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.368525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.043470Z digest=sha256:0bf7f23ee2e557d58b731dd4798fd6fcd6b32081bc9a431645c797e8c4faa75a

Observation c65daf9e-c0af-46ef-aa7d-b44bcf0bcbee · outbound

This paper cites Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...).

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Lemma D.2(Self-normalized bound for scalar-valued martingales).Consider random variables (vn|n∈N) adapted to the filtration (Hn :n= 0,1, ...)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.352347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.049084Z digest=sha256:98143beb3b58f14a992851b11c7353f19422113c2a7feccd73af3a7e81e81b1a

Observation 15ae85a1-92f2-4ddd-a953-0fecbb8f89a3 · outbound

This paper cites Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Table 16 shows the MLP architectures, i.e., GIRIL’s encoder and decoder and V AIL’s discriminator

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:01:05.419654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.028464Z digest=sha256:e2601e951715d741d93f2377f3b2e6bedd8e66dbb1c7b9e014fed1825d3343cd

Observation 3dca2e0c-e864-40b9-ab17-07e8e35015c4 · outbound

This paper cites Toward the fundamental limits of imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Toward the fundamental limits of imitation learning

Reference 1988

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.844783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.877390Z digest=sha256:1162b6b6da6c80e96d7bc570a924e5176237f4b36eed31f89390fefef6343d32

Observation ed4f5999-4c14-45f7-a129-809e339dc0d1 · outbound

This paper cites Relative entropy inverse reinforcement learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Relative entropy inverse reinforcement learning

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.149046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.757294Z digest=sha256:984b6b4c86a6b868ae740bc079231bc89443d57b7ae0c7f115ff31d7b6a841b0

Observation 5d5afa12-317f-4884-ad1f-fee39e977ae9 · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Learning structured output representation using deep conditional generative models

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.765510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.915792Z digest=sha256:4f679e27ec9c5639d5b05bfcb9f5d1bd3fe6f0abff8ef933a2dc57db502b853a

Observation b5aefac0-2d17-4260-a7a9-046d7ec9f654 · outbound

This paper cites and Bagnell, J.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration and Bagnell, J

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.586670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.978061Z digest=sha256:c7cb431959043e735de4770b300ca7886171d1c2d4a20001c2105eb989cb66cd

Observation 49dc83bb-418a-48f0-a380-a4f038286589 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Provably efficient reinforcement learning with linear function approximation

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.975912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.836637Z digest=sha256:a0228de192beed58a6bfa5ec276c4423b1dbfb4ee604d03bc3732901f348ff85

Observation d20c8eb2-84bc-45df-a5cb-92e43fa43df2 · outbound

This paper cites OpenAI Gym.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration OpenAI Gym

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.762311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.762311Z digest=sha256:ef2c578770391dfaffffe400841f6cba6d39a0e52a7e1230f9c3e4ced7922e4c

Observation ffb10c27-9f1e-4529-a1b7-1c24fc875d59 · outbound

This paper cites Proximal point imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal point imitation learning

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.731530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.925993Z digest=sha256:913932b372f21f2f60a185158dfce0e532c44e97013ae5c73adfe5d16e7322bc

Observation 35da6538-f151-47cf-8194-bc640d1ed319 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.893867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.893867Z digest=sha256:0a164c0a1fa9e0aa11da561ea445ee380120c80b9d4bfeb3b493e09b964228bb

Observation 9b272b49-a52a-4f0a-9472-499bc4ac3db2 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Curiosity-driven exploration by self-supervised prediction

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.895423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.864194Z digest=sha256:c8f7d691a5d08df1111466549e9e2f153bdf5a2b9042dff0e1ef2ca44942df6d

Observation e4ddb88d-b381-4354-b3c4-c767fd83b6c0 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Mujoco: A physics engine for model-based control

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.748282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.921365Z digest=sha256:89a038b2f74233916a92f7a5a4b63945dba2bfdaf368f2cb7c2a9a7bd1049de8

Observation 22552bf1-f381-4fa2-b776-3e639ded67e3 · outbound

This paper cites Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Extrapolating beyond sub- optimal demonstrations via inverse reinforcement learning from observations

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.133533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.768492Z digest=sha256:e5dcde7921b809d89659f9b6de89074396b3fe06874028713c18be6824ab3189

Observation 43d0ada8-eb40-4889-bc2d-e2d7d8a241d2 · outbound

This paper cites End-to- end driving via conditional imitation learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration End-to- end driving via conditional imitation learning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.088173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.800160Z digest=sha256:cbbe3cb76a19a939a96a0225854e14dfe3d20f6dae4e6810c2b4bc97b233b36c

Observation 8fc4b90e-23d8-42d3-8a74-998af7883bec · outbound

This paper cites On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.300489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.778839Z digest=sha256:eedcb9f4d7420f867c7b98bb7e1e8fe4ef6fae32917a15c6d2d6c96b3bbc19fe

Observation d5d1a19d-f8db-45e8-8aa7-21be4d6cd5cb · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Large-Scale Study of Curiosity-Driven Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.773534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.773534Z digest=sha256:fe3ef702f11cfda9616b87b9aa0aaaba1e7971003640737cfe92e3acc575e8d4

Observation e83f2d93-b30d-44a7-afce-4c234c4597cf · outbound

This paper cites Hybrid Inverse Reinforcement Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Hybrid Inverse Reinforcement Learning

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:01:05.226854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.882451Z digest=sha256:987aef269451e557b4e3887c0f94cdff5f8f8fceb326d700047854f4142e0be5

Observation 82a4e234-d90e-4be9-9f52-028324fb64e7 · outbound

This paper cites On Computation and Generalization of Generative Adversarial Imitation Learning.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration On Computation and Generalization of Generative Adversarial Imitation Learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.117953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.789475Z digest=sha256:a98a1cf67529b6850b2f81f7648e7edf36078ff39645a7576773fd8f56846874

Observation f1698227-9077-41f4-b4ab-59ce9bea3dab · outbound

This paper cites Imitation Learning from Imperfection: Theoretical Justifications and Algorithms.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Imitation Learning from Imperfection: Theoretical Justifications and Algorithms

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.943710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.845606Z digest=sha256:c84a9929abf9074da795c120099713e72d83fd327f00e8dc1cd157534544903d

Observation 47b4d814-94b0-4d36-972e-c16de17b9058 · outbound

This paper cites IQ-Learn: Inverse soft-Q Learning for Imitation.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration IQ-Learn: Inverse soft-Q Learning for Imitation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:06.025396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:04.818299Z digest=sha256:fa0e8fd89c842319e563d443130f8be4c0ea2ca959bd9a29378a8f822cb02318

Observation 8276749f-f8d1-469e-8c3e-e8d28e216c89 · outbound

This paper cites Exploration via Elliptical Episodic Bonuses.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration Exploration via Elliptical Episodic Bonuses

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:04.822802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:04.822802Z digest=sha256:110b689c8e54c02d07438f3dccf38d0628b4632616bb847c4e691e56e57039bb

Observation 009e47ab-f081-4445-8d6d-f6bf83f30407 · outbound

This paper cites For a fair comparison, we used an identical policy network for all methods.

Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration For a fair comparison, we used an identical policy network for all methods

Reference 5184

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:05.488424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:01:05.007817Z digest=sha256:464baed03ae757e8704b3e349f6546bc82f3e2f45bd8db6b312cd46ddf133d1e

Pith citing papers

No inbound Pith citation observations are available.