Pith. sign in

Paper Citation Record · LEDGER

RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2405.00254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00254 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:05.229990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:28:48.418933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 310957cd-3dfd-4981-9928-d4c190692c08 · inbound

MOSLIM:Align with diverse preferences in prompts through reward classification cites this paper.

MOSLIM:Align with diverse preferences in prompts through reward classification RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:05.229990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:05.229990Z digest=sha256:b76f867890d28ed35c8b8ee0c3a5c2ace349ccc52c5318ce8fc6d2b6902e284f

Observation a3775ccd-ef83-4694-be14-831abd0f9ccf · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:09.533133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:09.533133Z digest=sha256:d7dd385b46668b39cc4152574f65c2bbec0aa88ba834d88ad647a650e973ac83

Observation 88259934-527d-46dc-b40b-0a6c150812d2 · inbound

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory cites this paper.

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:31.735488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:04:31.735488Z digest=sha256:83cf1962b0596f51e61d091a890249b39670a41f5db62b412ec4645b5a9637e7

Observation 14f7efb8-29bc-47d7-8750-2cda7b46eed0 · inbound

Uncertainty Quantification for Ranking with Heterogeneous Preferences cites this paper.

Uncertainty Quantification for Ranking with Heterogeneous Preferences RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:13:14.351163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:13:14.351163Z digest=sha256:1a99fdf2e81d6f5df81e6a55442cd72fa5a41e22be3a3c76b54ff297042f9549

Observation 239942ba-700d-42ff-ba00-a3e876b9be0d · inbound

T-POP: Test-Time Personalization with Online Preference Feedback cites this paper.

T-POP: Test-Time Personalization with Online Preference Feedback RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:12.190762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:12.190762Z digest=sha256:57a5a3fdd95c21cc1a1a04e552d4f4c102f0bbc183c8c601abcde0c152de0950

Observation fae086ed-9e72-481f-b3f7-ab60576aacda · inbound

Collaborative and Efficient Fine-tuning: Leveraging Task Similarity cites this paper.

Collaborative and Efficient Fine-tuning: Leveraging Task Similarity RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:49:40.071442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:49:40.071442Z digest=sha256:3da26ddc7b12832b1082d626eb24ee03ef73eeb0a2b6713a0606fdcafbb8ef23

Observation b952dc43-dab5-4a02-ba32-3f648395960e · inbound

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration cites this paper.

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T10:27:28.560210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:27:28.560210Z digest=sha256:67702e0cd0f7a74a44b2baf45a0a45323466ab160f761dae762ba24ed37c78ef

Observation 7e3424c1-875a-4fa4-8a9c-49bc081ee610 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:54.490353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:30:14.693348Z digest=sha256:9498cbcd44864fc606075436b1cf9cb5e91ee5e4f163bc45ddf11feff8452ce3

Observation 9aa0016b-0d44-44a2-b1bb-81773a671c81 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:45.953490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:07:51.683581Z digest=sha256:fda21319811de69f9338d72ed8baf78392f40e6561ee91cc210ce8ae1076c147

Observation 6c2f6044-ea51-484e-b2e7-792df8212b65 · inbound

In-Context Reward Adaptation for Robust Preference Modeling cites this paper.

In-Context Reward Adaptation for Robust Preference Modeling RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.196063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:19:55.177440Z digest=sha256:2a46c6d5ff9709c25388ffa70c61a6fb3df7fdafb2d7e04bf851ef262c0204b9

Observation f6755436-d58a-473e-bcd9-f504a92a7e35 · inbound

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach cites this paper.

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:49.757731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:58:10.258001Z digest=sha256:f76882de8287bef663ad4ad79fb366d60f4486f6aa9e6b110d8184f95be011d9

Observation 6798f687-da7b-45f6-9998-02404fd32324 · inbound

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling cites this paper.

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.632083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T20:00:05.900814Z digest=sha256:dc0c649d4f0dbb66167d0145f4457e6f188431d496feb00504d428465338980f

Observation 93c26ac0-237b-4446-9be5-d9e50eaf5f3d · inbound

Hidden Consensus:Preference-Validity Compression in Human Feedback cites this paper.

Hidden Consensus:Preference-Validity Compression in Human Feedback RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:07:39.038708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:25:24.761133Z digest=sha256:014d884df875aed9650d2204e70468c3ad46636786687020b1a2b17a2bbbcb7d

Observation ccad351d-7b56-4eb5-8122-4fc3ec027d10 · inbound

CoPersona: Collaborative Persona Graphs for Robust LLM Personalization cites this paper.

CoPersona: Collaborative Persona Graphs for Robust LLM Personalization RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.420560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T18:20:53.930801Z digest=sha256:1053a7462320a6e41538c1f6b1c138dc2d4f88794ae25b2c84d107237f79f877

Observation 7d7103a2-8431-4882-8504-9b91fd8f6d1e · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-31T23:51:56.233532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:51:56.233532Z digest=sha256:2d7d003e48e1fd15eb33023a7cd75a298d8f8064aa2b132eef091ed29c34af4a