Pith. sign in

Paper Citation Record · LEDGER

Preference Transformer: Modeling Human Preferences using Transformers for RL

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2303.00957.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.00957 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:52:24.143019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.217026Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d11b751f-1c02-46ca-9c8e-d1027f6df6b0 · inbound

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment cites this paper.

Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:51:53.196903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:51:53.196903Z digest=sha256:14326f480a7b4ed7453fe37635c9f04548a30e0ef995497aabbc466ea3057750

Observation 2809a598-64be-4929-b590-d4fc542cf60b · inbound

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective cites this paper.

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:40:10.017433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:40:10.017433Z digest=sha256:7c835a171aaab5933c5592bd1766be77b453abd82d6076c21d3efed599806e01

Observation d193b40d-bb8a-4dc6-90b7-c540a4f1cc73 · inbound

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation cites this paper.

SimulPL: Aligning Human Preferences in Simultaneous Machine Translation Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:21:56.094516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:21:56.094516Z digest=sha256:0f7468332d3743e9d2664e84f051df0c8dec1931ce152a6d792020b80b9010a7

Observation 22b1ecfc-84e9-48c5-b1a7-520499d955fe · inbound

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations cites this paper.

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:52:24.143019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:52:24.143019Z digest=sha256:67b4ca5f6fff1072bfb0230ca7fc83e661ac569650a636706faa6984aac30eb8

Observation 1b95cad8-1d9e-4eea-b430-9a45b1b06e2e · inbound

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning cites this paper.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.589096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.589096Z digest=sha256:0ed84af301765a2d8f956971311b6945d7b523a18aac283ec82760875c198933

Observation d6570f42-3b80-4095-b1cc-01223c13b89b · inbound

Reward Models in Deep Reinforcement Learning: A Survey cites this paper.

Reward Models in Deep Reinforcement Learning: A Survey Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.639257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.639257Z digest=sha256:bbe15f0302b96ec92be9aecf841c2590e7f8683099e240d3d249d9c827b692d3

Observation 400ac5c5-b05d-4e79-8f8c-9218e0d3ccca · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:21.024330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:908f546d7cac3f1c29ea96a93f8b180b8bf515d05d0e31d5c17547b164f48c4c

Observation a91461f9-1259-4997-b09b-baf9131b7999 · inbound

MAPL: Multi-Objective Preference Learning for Robot Locomotion cites this paper.

MAPL: Multi-Objective Preference Learning for Robot Locomotion Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.218526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-25T21:11:00.753949Z digest=sha256:0ae6c347cca79ca06bc4f66dc33a99e17c19de80b5935ea9109c2a95653fca33

Observation 15fc3cd5-f288-4405-864a-c698ba0c5770 · inbound

SPLC: Social Preference Learning for Crowd Robot Navigation cites this paper.

SPLC: Social Preference Learning for Crowd Robot Navigation Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:18:06.215613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T12:10:03.492401Z digest=sha256:36eaf1d88db0b1c9600a55bbdc210d4d50fa2029cdebaa701189b57a2f4276f4

Observation 5c5e12a0-2ffe-4969-81fd-e524c9241314 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 227

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.065706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.065706Z digest=sha256:2eb3b6b98b7b6668a110fc96c1869412c0d7015f288c23bbab82337eaa843885

Observation 38abb7ee-237b-4218-993a-661c3d4087de · inbound

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling cites this paper.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.613860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.613860Z digest=sha256:6fe958d824549c6dcce994ba4e1cf655166b081274c7d770fe5dbb32ca581508