Pith. sign in

Paper Citation Record · LEDGER

The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2311.00168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.00168 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:44:51.027151Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 681a28bd-bd91-48d0-9441-54926090c9a0 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:53.473020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:392b96003e55e46bc338edf1c7c0eb20656cb7990fe401112922c0abd43b4826

Observation 6e4466c0-7a78-42d1-8a35-87b0fad540b6 · inbound

Drowning in Documents: Consequences of Scaling Reranker Inference cites this paper.

Drowning in Documents: Consequences of Scaling Reranker Inference The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:13:43.336223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:13:43.336223Z digest=sha256:2159012465e01593f4e053d6f0f353839cf8d3cc399d5f930f03ddb5e93a7e2c

Observation 18fe0b8c-1553-4484-875b-0d72e477049e · inbound

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling cites this paper.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.206702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.206702Z digest=sha256:11643b326b8d9998f68c52d1f90b43ca1cf948e4455b35827123de27be9f06db

Observation b5762376-55e6-4638-8078-b0f1ab623c98 · inbound

Establishing Reliability Metrics for Reward Models in Large Language Models cites this paper.

Establishing Reliability Metrics for Reward Models in Large Language Models The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:44:51.027151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:44:51.027151Z digest=sha256:40439d71956c0f4ab6bc1daf33859f75f355d95ed943a9651ed3cace0b6261c8

Observation 2d3a95c6-e250-4a0b-8ff2-59dfa125a0ef · inbound

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF cites this paper.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.230482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.230482Z digest=sha256:608a681a8a221072136ea33369a4bf0fefb7bc4ed0e1dc3c05925c20aabcfc21

Observation 92b21f53-dc94-49df-ad26-ea6491b33297 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.451656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:047efd7fc614cc9a0565145e5f09f381adfee612c5cadb66560cf34a78e981ce

Observation d37aecc7-4f51-4543-9e9a-8999ad36b06b · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.854732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:4e1a9ffabad17f85c59669512281467df0ad3ece022aa7e573ec66facd169adc

Observation e0b4ec5b-9979-40ce-b916-2db4a4c0f7ba · inbound

In-Context Reward Adaptation for Robust Preference Modeling cites this paper.

In-Context Reward Adaptation for Robust Preference Modeling The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.182550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:19:55.177440Z digest=sha256:ef1e12df99571ba03a19577913ebab47c28c520873dbe4d13376fca8d5fd2e56

Observation 53c5fe0a-59ab-4df1-88b6-f88cf772a7f3 · inbound

What Do People Actually Want From AI? Mapping Preference Plurality cites this paper.

What Do People Actually Want From AI? Mapping Preference Plurality The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:41:29.695819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T01:32:04.660400Z digest=sha256:fdddf12ca34d3ba179c478d2305d9358bca44c40b3730dc226da8443173b469c

Observation 65ac556c-c160-40ef-a9ca-dfe763c34ee8 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.492205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:0a3b204d1f2a8c22d9cf038b7f037ffede32fae5b277ed3c4fdd3f05c9c31255

Observation 82d42bf6-2cca-4297-ae8f-a32229b733be · inbound

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback cites this paper.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.540226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.540226Z digest=sha256:48aa358ed089c9ea3d089fa6beb383acc25825076479c1a6f9d033747ed00858