Pith. sign in

Paper Citation Record · LEDGER

Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2312.08358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.08358 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:49:10.543390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0531dae-9d5b-44ea-8ec1-4e8c62143004 · inbound

Test-Time Alignment via Hypothesis Reweighting cites this paper.

Test-Time Alignment via Hypothesis Reweighting Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.535424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T06:55:54.051821Z digest=sha256:24759a25f584df765f400e68b1d7d6e97d854782e3579de010992ec09132e42d

Observation 5d19d2f2-1c7b-4921-a9fd-08f0ea002ba9 · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.543390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.543390Z digest=sha256:3f45b5fbb816c1cf90667fa65f46cdec2c205f169ebd01f9c95f66fd7f26964a

Observation 699e2fc7-0744-4fe0-9bf4-189127a6bc72 · inbound

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory cites this paper.

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:32.028901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:04:32.028901Z digest=sha256:81b5e02a3b19aa379471576344cff47c5a937554b6a9d59dc08aa8401811ed28

Observation 129db947-a33a-4513-93e2-67f981934587 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.197977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.197977Z digest=sha256:378b3d4df2c8b60df8f258d89a46e070665c909c4c4dce164cd706a6207472f1

Observation e20a2f2a-12a6-4e2e-8637-5519f87425ea · inbound

Active Query Selection for Crowd-Based Reinforcement Learning cites this paper.

Active Query Selection for Crowd-Based Reinforcement Learning Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:02:26.007783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:02:26.007783Z digest=sha256:62196b21846153692786c7e3c8a223ec7424886c421e15c221fb934ae634d74b

Observation 1abf0f85-8f73-4ef6-801e-8190f4448a73 · inbound

RLHF May Not Reflect Genuine Preferences cites this paper.

RLHF May Not Reflect Genuine Preferences Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.322374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:39:08.880486Z digest=sha256:404dbdbe1bdd87cffc26ea110f607648e2b95471e6c49c7a9cb575c0ce445098

Observation 4fc5fe13-7443-4442-8716-83ee117809d9 · inbound

Efficient Personalization of Generative User Interfaces cites this paper.

Efficient Personalization of Generative User Interfaces Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:00:57.559506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:52:34.799158Z digest=sha256:1392c460748271b807ad37bcdfc9d054ac5b3dade017751b9a11b7c76de07ce6

Observation 06cf540a-7373-4691-b261-15b869a81079 · inbound

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem cites this paper.

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.696412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T23:02:34.564375Z digest=sha256:dda84d27828e5056337d9da4b94087e8b54b2825d0cf1a237bb806c6457e2f9c

Observation a4e556fa-f511-4cbe-b4cd-41ca01a10e4c · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 186

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T16:51:06.005505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:5df9b1001885e44f51573ec9a6a2e7f8fd61cbab9312fdcf2e6d69e4c9188758

Observation 7c1484f3-54f5-4e27-b71b-48038f27c286 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.259581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:8c72b03a4802a13546fe5c67c5ed92be2b53fc3d100ac78d3ef4bb054c0c6dfa