Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment

As of 19 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 0 inbound Pith citation observations for arXiv:2601.22823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22823 v2

Coverage vector

measured 4 of 4 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:28:01.767455Z

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

4 of 4 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 239c8204-ed52-42f1-ab81-856ad4e8e32f · outbound

This paper cites write newline.

Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:28:01.282100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:28:01.282100Z digest=sha256:b92ed09bb5385014ed475dfe41e23ecb081c19d20664fa5f880e67f5706d2173

Observation 839e90d2-cfda-4823-b129-32bac2686205 · outbound

This paper cites @esa (Ref.

Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment @esa (Ref

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:28:01.441524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:28:01.441524Z digest=sha256:c501514715f6db8b6c3e9d22190e542008526a227a90fd94d58e8f3e0cdb71b2

Observation 97231f5c-9352-4909-b580-905d07dbb7a1 · outbound

This paper cites an unresolved cited work.

Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:28:01.589621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:28:01.589621Z digest=sha256:b7e0fdbec9a8a88bd8bf6f89974b2948b466efa264ba242a941f0687c6186012

Observation 3664f6e3-c1a3-49a5-8e6b-83c0ec46a218 · outbound

This paper cites the vector of an unsupervised learned trajectory encoder.

Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment the vector of an unsupervised learned trajectory encoder

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:28:01.767455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:28:01.767455Z digest=sha256:ead94e0de6ace1f1d22b8a40b0b74aed6522a99f45d57e3b503fdd6d7ce14df3

Pith citing papers

No inbound Pith citation observations are available.