Pith. sign in

Paper Citation Record · LEDGER

Models of human preference for learning reward functions

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2206.02231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.02231 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:58:38.566308Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:38:04.667531Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f398cc95-f2a7-411e-910d-76a946347aab · inbound

Solving the Inverse Alignment Problem for Efficient RLHF cites this paper.

Solving the Inverse Alignment Problem for Efficient RLHF Models of human preference for learning reward functions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:14.429741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:55:14.429741Z digest=sha256:7593a1370ecb63f5092d33ad7901c2a78350257f4ca20bb20075f20d324e3540

Observation b019230c-ff8d-4849-aafb-1951c269198e · inbound

Policy-labeled Preference Learning: Is Preference Enough for RLHF? cites this paper.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Models of human preference for learning reward functions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.566308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.566308Z digest=sha256:f8a85d859e701a67c5825379dfa1ab95b890aeb317954bb3785529c105423a2e

Observation 32592992-01cb-4497-a287-6d56e90f44a0 · inbound

DMRL: Data- and Model-aware Reward Learning for Data Extraction cites this paper.

DMRL: Data- and Model-aware Reward Learning for Data Extraction Models of human preference for learning reward functions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:39:04.549420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:39:04.549420Z digest=sha256:8288c1e47e9660c654f0249ebe44a00f1a89b75f542413755a7fbe87d07f0aae

Observation baff5cc1-6d0e-4bf8-adc7-3bcd9acfebef · inbound

Reward Models in Deep Reinforcement Learning: A Survey cites this paper.

Reward Models in Deep Reinforcement Learning: A Survey Models of human preference for learning reward functions

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.656340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.656340Z digest=sha256:b74ad30a4d4d6a0769715f3644560853f807f71ca0dd454f42112145f5143003

Observation b05217fc-f97d-452d-9eea-6474cbc25f63 · inbound

Misalignment from Treating Means as Ends cites this paper.

Misalignment from Treating Means as Ends Models of human preference for learning reward functions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.199300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.199300Z digest=sha256:fd72f0dc9aded911e1c04fdb499e9ce767282e99749a4c9650aa038bbe785f70

Observation 07537b37-3f91-4156-8081-dfd99ef58de5 · inbound

PrefPalette: Personalized Preference Modeling with Latent Attributes cites this paper.

PrefPalette: Personalized Preference Modeling with Latent Attributes Models of human preference for learning reward functions

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T16:27:57.235072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:27:57.235072Z digest=sha256:7f5bb3f06aebc58ee3af12a2b901ade7e90a8441739ea001d3a8855976b37600

Observation 4e95789a-6a4c-468d-bd98-2e455c7d88da · inbound

Mitigating Cognitive Bias in RLHF by Altering Rationality cites this paper.

Mitigating Cognitive Bias in RLHF by Altering Rationality Models of human preference for learning reward functions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:57.056054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T01:02:20.179748Z digest=sha256:a2c1d35bcc31c2f8c56e083c378422a1c2350626e2aa9fe5b82dff78ed3fc592

Observation cc8c8ecb-6ff7-48ba-83b3-8c666bd9948a · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Models of human preference for learning reward functions

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.249499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:b92e19c18e2254c9c59a8958c819eacc71e7bd4df0a2ea64b6d645b423e2d991

Observation 1d75e64d-a71d-4dca-bcf1-f1feddab9c6b · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Models of human preference for learning reward functions

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:06.432792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:9e47afccea7296f708daffbebaa4ae5d15807da4c3be6ad603611874d9cafe01

Observation 45a86250-9081-46fc-b706-3fb36278c0c5 · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Models of human preference for learning reward functions

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:31:16.901456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:87f0d99ec31bd26fe9771a06752a872e4bf89e617c8c7020c1a28f7b86c3591e

Observation 4ab85991-56ed-4bba-bfd0-31af34265e63 · inbound

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation cites this paper.

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation Models of human preference for learning reward functions

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.669002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:29:43.030058Z digest=sha256:d6b92a0f0d1bfb94ee45c7cf081a0996622ef42ccc787b5376267c5a3ec9ea18

Observation e3b2b4b0-5c97-47fa-a9ab-0c4d8ab12a8b · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Models of human preference for learning reward functions

Reference 250

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.139912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.139912Z digest=sha256:76506ad59ca81fd6dd50b494382050ce3fd7a991fe4987e05af748586bf40a8a

Observation 34a6412f-ff93-4201-886c-6617e19fa7e0 · inbound

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling cites this paper.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Models of human preference for learning reward functions

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:45.573168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:45.573168Z digest=sha256:28dd8411b24063a7ca675c0e45bb1a233569e1990c37d2c390433a15640615e5