Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2501.18873.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7c957e81-5093-48bc-8e18-58bd8cce6e4b · outbound
Best Policy Learning from Trajectory Preference Feedback URLhttp://www.jstor.org/ stable/2334029
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9716bf0d-1090-4ba4-bbcc-d596d3baf637 · outbound
Best Policy Learning from Trajectory Preference Feedback Bridging Imitation and Online Reinforcement Learning: An Optimistic Tale
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b7bc4cc-0b50-4c67-af02-fbefef720f53 · outbound
Best Policy Learning from Trajectory Preference Feedback Introduction to the non-asymptotic analysis of random matrices
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d612f9c0-d448-4efa-8b82-0411b37fac82 · outbound
Best Policy Learning from Trajectory Preference Feedback Yes, please see Section 2
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c05d7186-4773-4fd8-bf18-3412c3960122 · outbound
Best Policy Learning from Trajectory Preference Feedback Yes, please see Sections 3 and 4, and Appendix A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef10d24e-7e89-475a-8eff-424e9868ba93 · outbound
Best Policy Learning from Trajectory Preference Feedback Yes, code will be released later
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08aae76b-2c3a-420e-96e3-23f5f470d362 · outbound
Best Policy Learning from Trajectory Preference Feedback Yes, please see Appendix A
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f141bf8e-9d16-4764-82c8-be0a6d35d179 · outbound
Best Policy Learning from Trajectory Preference Feedback xp1´xq.fis a concave function. We have for anyiP t0,1u, Prpπpiq k ‰π ‹q “E
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 22f8b4bc-eb67-4c6b-97ec-17f90f2be540 · outbound
Best Policy Learning from Trajectory Preference Feedback Notice that the eachchpsq is the difference of two binomial random variablesb1 „BinpN, 1 ´γ β,λ,N q andb 2 „BinpN, γ β,λ,N q
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d16f3d24-4952-4cfc-a931-f76e293903b1 · outbound
Best Policy Learning from Trajectory Preference Feedback argmax θ,ϑ,η Prpθ, ϑ, η|D kq
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b140fbc3-57ef-40ee-8683-848dd13eb80e · outbound
Best Policy Learning from Trajectory Preference Feedback Similar idea has been proposed to estimate the expertise level in imitation learning Beliaev et al
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c844c675-38e8-4251-b014-773d062a375d · outbound
Best Policy Learning from Trajectory Preference Feedback 1 diam ` Ft|xt ˘ ďα`Cpd^Tq `2δ T ? dT , whereδ T “max 1ďtďT diam ` Ft|x1:t ˘ andd“dim E pF, αq. Lemma B.8.If pβt ě0|tPNq is a nondecreasing sequence andFt :“
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.