Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:39:00.790790Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.29617.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:39:00.790790Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a32e3473-0170-4fed-a666-65629cd8e3cc · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e01af168-0c2f-4314-a5a7-fd521672d600 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce16b7a1-9f83-4eed-8170-318139427623 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning 57Appendix table of contents Part II Additional Results We collect here additional results omitted from the main text
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a047ee89-3ac2-4571-a270-4c4421cad7f2 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning It is strictly suboptimal, with J πb Mb = 1/2 and supπ J π Mb = 1
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb40963-218d-4bd4-afc8-8ac14b8260ca · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning If {πb :b∈ Bn} ⊆Π, then, for everyε∈(0,1/2), log2 Nε(Π, dΠ)≥2 n −n−1
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55af81c0-6ca3-4570-83d8-cd69e9d956f5 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning , QK ∈ Qand coefficients (wk)K k=1 defining the linear combination LC((Qk)K k=1) = PK k=1 wkQk
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26c9ae6-8355-4920-9dd5-fe5a0e3608a1 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning It follows that, with probability 1/2, the next state is x−
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af1816b-a5f5-4da2-a999-8b230c982408 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning The resulting algorithm is in Algorithm 3
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e96416-ff2d-4fff-ab0e-f565bbaae6f4 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da324c86-8952-4077-b147-cd90ff6b1c66 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c088e1c1-8d8e-4e26-84d1-3effefb409d1 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning HX h=1 X x∈X dπE h (x) D QπI h (x,·), πE,h(· |x)−πIh h (· |x) E# . Usingd πE h =d h + (1−α)(d πE h −d πout h ), the right-hand side is equal toT 1 +T 2, where we define T1 :=E I
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eee1d9d-00ca-4ed2-a928-982f3d09fb48 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning value-based IL
Reference 1991
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fbca4f4-5c12-46cc-a076-64de5c06e280 · outbound
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning A Theory of Learning with Autoregressive Chain of Thought
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.