Pith. sign in

Paper Citation Record · LEDGER

Reinforcement and Imitation Learning via Interactive No-Regret Learning

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:1406.5979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1406.5979 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:30:49.157907Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T08:17:45.248813Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a53853e-3cbc-4f46-8be1-ca49315e7df5 · inbound

Learning Belief Representations for Imitation Learning in POMDPs cites this paper.

Learning Belief Representations for Imitation Learning in POMDPs Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-25T17:56:06.648842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T17:54:12.840437Z digest=sha256:358351ba41b488aad1edfc87f860bb7c80240e7963cc7baecbcb94898f94112f

Observation 43814708-aa17-4c39-bdbc-ecd51f5ff171 · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:42:19.101281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:49f32cdd2b26e47475e59dc36b41d50f60c9f3a039793438704ae0fbfdce668d

Observation 8b12cbb0-c742-479d-ae10-1566c49eac71 · inbound

State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning cites this paper.

State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:24:18.185166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T18:22:02.372554Z digest=sha256:0974e27c43a8c176888faf77c1e46357f108f5fbded1b9d8b8be9f8d74730a2c

Observation b325bd8a-c583-4862-85b8-638755cfab4a · inbound

State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning cites this paper.

State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:30:49.157907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:30:49.157907Z digest=sha256:35019f638caa1269e20c064418e7a4e16f8d5353f5131d4b2e61197eb45c529e

Observation 6e09e307-c4c6-4f66-808a-f4b7891234c6 · inbound

BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving cites this paper.

BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:16:09.078118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:03:02.833156Z digest=sha256:92d31418c6f37533d577fc0510c3bf63416144aba01293a16570d47d32a0e401

Observation 42dd62c3-8446-439d-8c22-aed3812b466b · inbound

Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents cites this paper.

Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:05:25.050088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:02:45.029951Z digest=sha256:ac598e335da7868cf741a8718af493fbe42cd51a8d15e69f2a0fda541d417ffd

Observation 22e50df1-5037-4ef0-8ab1-927df1d955fe · inbound

Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents cites this paper.

Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T20:36:59.298711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:36:59.298711Z digest=sha256:dc87f2a6bf065cc65e9a58ea02e6beddf592eb48bd1db33f318dca851c575861

Observation 9d5f0629-4a9c-41d1-85c7-20e3b10681a8 · inbound

The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation cites this paper.

The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:08:59.125288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:07:49.043189Z digest=sha256:0b2e44c20f823da5e349d27448d67750d672bdbaf001828507a94a8d53562457

Observation c4838a9d-3bb8-4cb4-b8ff-e2e84a43988c · inbound

Provable imitation learning for control of instability in partially-observed Vlasov--Poisson equations cites this paper.

Provable imitation learning for control of instability in partially-observed Vlasov--Poisson equations Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:11:05.274072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T16:32:45.224829Z digest=sha256:acf32577176ae69a360f98ddced8f2bee315824473893bb25c28f63ea1aa7583

Observation 3bf08f22-8d83-49dd-b4ca-e8470761cfcc · inbound

Revisiting DAgger in the Era of LLM-Agents cites this paper.

Revisiting DAgger in the Era of LLM-Agents Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:57:53.264138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:56:06.762156Z digest=sha256:77f35f3e5806d5951b7bd3dc5f3ac44b9ae8c50bb31d7471ea039298a23e758f

Observation 748ecd4f-f159-4602-a947-a3c939abad0d · inbound

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach cites this paper.

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:02:49.760079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T23:58:10.258001Z digest=sha256:48e6f86dcde4a007f107740e9e5c03519fb79622cc2f945df16528effc69b237

Observation abd24304-83cf-4192-80fa-107bc7b467b7 · inbound

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation cites this paper.

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T08:17:45.250472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T10:53:27.839408Z digest=sha256:c4877bcba03fc42ac83a36f45ee900f8cc28fae4b8ae9bc6d7e3720ba1e84171

Observation 3a4556c6-937b-4a99-b035-2d7504350ce5 · inbound

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon cites this paper.

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:44:21.694949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:41:39.266071Z digest=sha256:038a22ed39801b01d5751eaf53ef5fb3c7396f3b057c20c354c808807f475b04

Observation 344e9a74-af39-4882-8b3d-4db7ce7e542f · inbound

Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback cites this paper.

Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:45:39.360449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T06:21:48.752660Z digest=sha256:4876b8cf2ba26c57813416c2ed7b2192e142afa46ccf35bd9b2ecbd951af9a4d

Observation a7b4f5be-d21c-4bb6-9bb3-2a82f49e70c9 · inbound

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training cites this paper.

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T17:03:13.686066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:03:13.686066Z digest=sha256:a750c8030aae2a30607d56be28624cc395cc62d8b67407da2f6b934b27723a75

Observation d8c4c4d5-f76d-4998-a8ce-c3d174e4721d · inbound

CAST: Game Solvers as Turn-Level Teachers for LLM Agents cites this paper.

CAST: Game Solvers as Turn-Level Teachers for LLM Agents Reinforcement and Imitation Learning via Interactive No-Regret Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T02:58:29.334255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:58:29.334255Z digest=sha256:53cae39012a97995e5de18fe43df76cad7147e47f3931de80aa7627713819ebd