Pith. sign in

Paper Citation Record · LEDGER

Best Policy Learning from Trajectory Preference Feedback

As of 12 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2501.18873.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18873 v4

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c957e81-5093-48bc-8e18-58bd8cce6e4b · outbound

This paper cites URLhttp://www.jstor.org/ stable/2334029.

Best Policy Learning from Trajectory Preference Feedback URLhttp://www.jstor.org/ stable/2334029

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:57:33.842172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:c0723eabcb43cd2ffcab9e0d462aac0dcfcfa5d170fa25d2b2261e952281698b

Observation 9716bf0d-1090-4ba4-bbcc-d596d3baf637 · outbound

This paper cites Bridging Imitation and Online Reinforcement Learning: An Optimistic Tale.

Best Policy Learning from Trajectory Preference Feedback Bridging Imitation and Online Reinforcement Learning: An Optimistic Tale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:57:33.833975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:8221a7d47391fa021eb2e8e48175dd0d5bc2b54952a29e8f778b7c017b707e5e

Observation 3b7bc4cc-0b50-4c67-af02-fbefef720f53 · outbound

This paper cites Introduction to the non-asymptotic analysis of random matrices.

Best Policy Learning from Trajectory Preference Feedback Introduction to the non-asymptotic analysis of random matrices

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T04:57:33.791191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:15e286726460a602a8476893da68635aeabc08deab43f9a965422b097d6a5267

Observation d612f9c0-d448-4efa-8b82-0411b37fac82 · outbound

This paper cites Yes, please see Section 2.

Best Policy Learning from Trajectory Preference Feedback Yes, please see Section 2

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.012250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:f6e9385be52ef486213d717baaf41cd798a6666b1eebe9d35756ab57c4176b30

Observation c05d7186-4773-4fd8-bf18-3412c3960122 · outbound

This paper cites Yes, please see Sections 3 and 4, and Appendix A.

Best Policy Learning from Trajectory Preference Feedback Yes, please see Sections 3 and 4, and Appendix A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.003093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:ced475c89e1b3be251b6b3e25a4a9783a8a548a83885ce51c9dc2d592982b7c8

Observation ef10d24e-7e89-475a-8eff-424e9868ba93 · outbound

This paper cites Yes, code will be released later.

Best Policy Learning from Trajectory Preference Feedback Yes, code will be released later

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.009197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:c3da459546b1f2bad13c5973bbfc118b5c45742848268b16bf45ca93680e82dd

Observation 08aae76b-2c3a-420e-96e3-23f5f470d362 · outbound

This paper cites Yes, please see Appendix A.

Best Policy Learning from Trajectory Preference Feedback Yes, please see Appendix A

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.025666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:a1d43e809ebb4d68a29ce3ad79d4cb985a6c7086b0d317411404a653d5b52dee

Observation f141bf8e-9d16-4764-82c8-be0a6d35d179 · outbound

This paper cites xp1´xq.fis a concave function. We have for anyiP t0,1u, Prpπpiq k ‰π ‹q “E.

Best Policy Learning from Trajectory Preference Feedback xp1´xq.fis a concave function. We have for anyiP t0,1u, Prpπpiq k ‰π ‹q “E

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.015513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:767ca3eb0e3de2b9d616895647cb3d3e43bb043e95c01777743ed29cf1cc001f

Observation 22f8b4bc-eb67-4c6b-97ec-17f90f2be540 · outbound

This paper cites Notice that the eachchpsq is the difference of two binomial random variablesb1 „BinpN, 1 ´γ β,λ,N q andb 2 „BinpN, γ β,λ,N q.

Best Policy Learning from Trajectory Preference Feedback Notice that the eachchpsq is the difference of two binomial random variablesb1 „BinpN, 1 ´γ β,λ,N q andb 2 „BinpN, γ β,λ,N q

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.018840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:e84858a73a210330a6508ed9c42dd497d8db5aede3685ed5d6af7ced5ab95c3e

Observation d16f3d24-4952-4cfc-a931-f76e293903b1 · outbound

This paper cites argmax θ,ϑ,η Prpθ, ϑ, η|D kq.

Best Policy Learning from Trajectory Preference Feedback argmax θ,ϑ,η Prpθ, ϑ, η|D kq

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.006114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:3aea06beb0fe9065f613d7cba5f2e665d464cab61e95b5762154523c5c6528c0

Observation b140fbc3-57ef-40ee-8683-848dd13eb80e · outbound

This paper cites Similar idea has been proposed to estimate the expertise level in imitation learning Beliaev et al.

Best Policy Learning from Trajectory Preference Feedback Similar idea has been proposed to estimate the expertise level in imitation learning Beliaev et al

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T04:57:34.028926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:a166af32c804cd036601e202922421bddd360fff9fcd6c1aedccfc450ad42a99

Observation c844c675-38e8-4251-b014-773d062a375d · outbound

This paper cites 1 diam ` Ft|xt ˘ ďα`Cpd^Tq `2δ T ? dT , whereδ T “max 1ďtďT diam ` Ft|x1:t ˘ andd“dim E pF, αq. Lemma B.8.If pβt ě0|tPNq is a nondecreasing sequence andFt :“.

Best Policy Learning from Trajectory Preference Feedback 1 diam ` Ft|xt ˘ ďα`Cpd^Tq `2δ T ? dT , whereδ T “max 1ďtďT diam ` Ft|x1:t ˘ andd“dim E pF, αq. Lemma B.8.If pβt ě0|tPNq is a nondecreasing sequence andFt :“

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-05-23T04:57:34.022420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T04:57:02.705703Z digest=sha256:3a8fe403a1113234e5b054669381875511c09f87cf8ad83b28c684b533e8984a

Pith citing papers

No inbound Pith citation observations are available.