Pith. sign in

Paper Citation Record · LEDGER

Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2411.04991.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04991 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:04.364900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:32:35.329195Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49292bd6-b8fe-4ec8-9aa2-a8a55676064d · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.667863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:4fa857f320ba2a2c35eb6ffbfff77d076aa131f7c79e79a7d88e6bf9133813f2

Observation a13a5ca8-7485-485b-9904-bfad94f57cc4 · inbound

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models cites this paper.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 1027

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.364900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.364900Z digest=sha256:eb4f28d722848c998d489a35a53657316c5c2e659dc208a408b8e8529e36e17b

Observation 1b8a5cf1-704d-4ee7-9611-ac9b10123401 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.213454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.213454Z digest=sha256:5018949a37bc1f715130c256b01ddfeb3feeb0daaa18dc95ff1da7e0739e480b

Observation faaa401d-b7f0-4ec2-b1aa-a54f8cd4e031 · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.374599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.374599Z digest=sha256:20ce2bde27a57f2d27cbb6b39658e9f6bfdff35637489fff857be5e31b90c9a1

Observation 4e368313-1aaa-4300-b279-2afba233b418 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:54.518269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:30:14.693348Z digest=sha256:1698b0421cd90973faa5edd060b09fc5aeb7c5803c907ee14b56fe0470740d31

Observation 4cf75416-2bca-4a70-b028-30d595d534e8 · inbound

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences cites this paper.

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.330785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T19:03:47.751245Z digest=sha256:882450d8ede4b721a120ba0c5e760988ad60f482013889427eaaa46408081027

Observation 59dad6fc-750f-4375-a283-f13f9e6a4c58 · inbound

From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation cites this paper.

From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:02:53.068235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:02:53.068235Z digest=sha256:1c5c0f3bfddce95867ccbda36f9797d7e576b523fe6e63e790c00bcf3c0f9b49