Pith. sign in

Paper Citation Record · LEDGER

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.29617.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29617 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:39:00.790790Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a32e3473-0170-4fed-a666-65629cd8e3cc · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.806437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.806437Z digest=sha256:447204bcf5a722af9f4a45e02e9f59070e0ff179237763a23eb48419c81092b9

Observation e01af168-0c2f-4314-a5a7-fd521672d600 · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.858491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.858491Z digest=sha256:a816954141c7c1b53d32649ff10515ca00967809d5bb1269f9f98861a646bed4

Observation ce16b7a1-9f83-4eed-8170-318139427623 · outbound

This paper cites 57Appendix table of contents Part II Additional Results We collect here additional results omitted from the main text.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning 57Appendix table of contents Part II Additional Results We collect here additional results omitted from the main text

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.355093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.355093Z digest=sha256:203a94ed82cba081b44d39b6f61628c0402822c38dd3d24e9a75b32696cc1721

Observation a047ee89-3ac2-4571-a270-4c4421cad7f2 · outbound

This paper cites It is strictly suboptimal, with J πb Mb = 1/2 and supπ J π Mb = 1.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning It is strictly suboptimal, with J πb Mb = 1/2 and supπ J π Mb = 1

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.960597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.960597Z digest=sha256:72b661548680b2764f05cccac1c4d9b5e901ded80ec7e8ebf4c6956dfa74446e

Observation 4cb40963-218d-4bd4-afc8-8ac14b8260ca · outbound

This paper cites If {πb :b∈ Bn} ⊆Π, then, for everyε∈(0,1/2), log2 Nε(Π, dΠ)≥2 n −n−1.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning If {πb :b∈ Bn} ⊆Π, then, for everyε∈(0,1/2), log2 Nε(Π, dΠ)≥2 n −n−1

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.090686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.090686Z digest=sha256:9712b12425f62ff922f097517bccb96938cd26f22d4592979501af661f344c06

Observation 55af81c0-6ca3-4570-83d8-cd69e9d956f5 · outbound

This paper cites , QK ∈ Qand coefficients (wk)K k=1 defining the linear combination LC((Qk)K k=1) = PK k=1 wkQk.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning , QK ∈ Qand coefficients (wk)K k=1 defining the linear combination LC((Qk)K k=1) = PK k=1 wkQk

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.157760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.157760Z digest=sha256:0b987b0f343b082de2428a87dc0bb1604ecda1e45205ac84ee36a1250ec556d3

Observation d26c9ae6-8355-4920-9dd5-fe5a0e3608a1 · outbound

This paper cites It follows that, with probability 1/2, the next state is x−.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning It follows that, with probability 1/2, the next state is x−

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.254929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.254929Z digest=sha256:aef8538354ce351ccd99d725f12e4669899571dcfbaef3ade18354050a0c3fc0

Observation 4af1816b-a5f5-4da2-a999-8b230c982408 · outbound

This paper cites The resulting algorithm is in Algorithm 3.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning The resulting algorithm is in Algorithm 3

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.433896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.433896Z digest=sha256:ffa3a81f203c0ef2710bf79c261f1a5fcbfa064e6d8ba224ddaee551e9d728f8

Observation 32e96416-ff2d-4fff-ab0e-f565bbaae6f4 · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.596173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.596173Z digest=sha256:c4d66b3ddcecff558b42c4da98302f3dba3e75eb09ee55c982524d436d7fd27a

Observation da324c86-8952-4077-b147-cd90ff6b1c66 · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.775798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.775798Z digest=sha256:ef10a9ca01b04ab9339190b75469b5634912ddc0b09e6f40cb4b54469aaf72fa

Observation c088e1c1-8d8e-4e26-84d1-3effefb409d1 · outbound

This paper cites HX h=1 X x∈X dπE h (x) D QπI h (x,·), πE,h(· |x)−πIh h (· |x) E# . Usingd πE h =d h + (1−α)(d πE h −d πout h ), the right-hand side is equal toT 1 +T 2, where we define T1 :=E I.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning HX h=1 X x∈X dπE h (x) D QπI h (x,·), πE,h(· |x)−πIh h (· |x) E# . Usingd πE h =d h + (1−α)(d πE h −d πout h ), the right-hand side is equal toT 1 +T 2, where we define T1 :=E I

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.790790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.790790Z digest=sha256:092586c066c6bacf4a9fe93a1e1f8b7e80a0d1b1485e6420787b5ede40b871fa

Observation 4eee1d9d-00ca-4ed2-a928-982f3d09fb48 · outbound

This paper cites value-based IL.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning value-based IL

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.609784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.609784Z digest=sha256:109892ad19684c8d491f98704bdbf79e2960abb46a276b56e6effd0258163c83

Observation 7fbca4f4-5c12-46cc-a076-64de5c06e280 · outbound

This paper cites A Theory of Learning with Autoregressive Chain of Thought.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning A Theory of Learning with Autoregressive Chain of Thought

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.481636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.481636Z digest=sha256:f2ba9a4ac995a66f777da91b22b48dd8c475d959a449535af1b2bf9e9d5503ff

Pith citing papers

No inbound Pith citation observations are available.