Pith. sign in

Paper Citation Record · LEDGER

Procedural Fairness Failures in RLHF from Preference Averaging

As of 18 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2608.10126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10126 v1

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:38.248174Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c92d9153-955f-4409-8c5f-ff86cac16482 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Procedural Fairness Failures in RLHF from Preference Averaging Fine-Tuning Language Models from Human Preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.192422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.192422Z digest=sha256:ba23efa900d27555a38bb3f0e12b933fc1e96843b44fbff4260e5967ad2b4a13

Observation d21df635-6ffb-4a7e-96ce-dc14de00ba1c · outbound

This paper cites Training language models to follow instructions with human feedback.

Procedural Fairness Failures in RLHF from Preference Averaging Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.198111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.198111Z digest=sha256:333be9d3bba4bdc8dfd05080e7bcae462937277b3ee8be6c994f360e2dd1d07f

Observation e3766b2a-9706-466f-a019-2b51725111a8 · outbound

This paper cites $\textit{Ab initio}$ dynamical mean-field theory with natural orbitals renormalization group impurity solver: Formalism and applications.

Procedural Fairness Failures in RLHF from Preference Averaging $\textit{Ab initio}$ dynamical mean-field theory with natural orbitals renormalization group impurity solver: Formalism and applications

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T04:15:38.406721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-14T04:15:38.207974Z digest=sha256:1c3bda22c78f5984cd2a54826f92a5e66b9fee930e5ec3a9752ec5de3fc2b198

Observation ec071893-9b6d-49e0-9113-6bd73884dc47 · outbound

This paper cites Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes.

Procedural Fairness Failures in RLHF from Preference Averaging Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.214421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.214421Z digest=sha256:cd6cc3ffbb8d5d93b4be9a2d0f138ae4013b0bb52e4bf7d62a675ded82567c1c

Observation 006d9a92-e3d2-4474-ae74-e729b55f641a · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Procedural Fairness Failures in RLHF from Preference Averaging Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.222855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.222855Z digest=sha256:3325d162ef64cdb55f30c9b4ec52dbb2108819fd789e352a1ce59dfb29e64633

Observation e19d68a7-45bd-4d12-a5aa-d51ef1e1a873 · outbound

This paper cites A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models.

Procedural Fairness Failures in RLHF from Preference Averaging A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.236029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.236029Z digest=sha256:e351ba99743803e34b395444aaf0b5c2e5606c9ef0ff7ff7239cc77d1bf610f6

Observation ca990641-9310-4004-a49f-2cf44770d3a6 · outbound

This paper cites an unresolved cited work.

Procedural Fairness Failures in RLHF from Preference Averaging Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:15:38.473629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-14T04:15:38.242008Z digest=sha256:5840a24fd57d26e553d1764c521de6576aa70445d4a3ff7b40aefed7183c8d86

Observation 00a91046-ce5b-4259-9251-58c85057b1f6 · outbound

This paper cites Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback.

Procedural Fairness Failures in RLHF from Preference Averaging Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.248174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.248174Z digest=sha256:0aff34bb123c734cedbbcb3a8f7aee21cde4a9368cab2acab0c8de88f169483f

Pith citing papers

No inbound Pith citation observations are available.