Pith. sign in

Paper Citation Record · LEDGER

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization

As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2412.03822.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03822 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:06:51.356046Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:32:04.660400Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 99e1cee2-2870-4e1b-8b6b-e677a898dcd2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.044240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.044240Z digest=sha256:a2d7d675710ad82a0dee8929e4c4bce09229a19bcc5404cb05fe0403ebc27c3c

Observation cfa3e659-91c5-40bb-8b5b-5b9a6682aea9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.050083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.050083Z digest=sha256:bb02a2c4e5d27033d37a81f855b654150a80034acd31ffa0bcf676e7997b6396

Observation db57cba3-db7b-466e-b44f-8a45c1b5ce2f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.056099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.056099Z digest=sha256:8c848470646c176599d39f23f74a44e473949fe8e5a636cb1ade09877f1fe286

Observation 2d6f3939-1a8e-41c5-941e-743146953699 · outbound

This paper cites GPT-4 Technical Report.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.061791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.061791Z digest=sha256:bd42daab6c16325338d1df449482f9fd45cda8fe56b963b8d422ce349aee07f0

Observation 22760f65-963c-4a7c-98d9-e29b8cac525e · outbound

This paper cites How Far Are We From AGI: Are LLMs All We Need?.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization How Far Are We From AGI: Are LLMs All We Need?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.067338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.067338Z digest=sha256:8bbec1c057856a12c288966f3d73b69546bb43205543499f6a14773fe4d3a453

Observation 97a45ed3-b108-43d3-b22c-bbd68b0d9935 · outbound

This paper cites Learning to summarize with human feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Learning to summarize with human feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.073149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.073149Z digest=sha256:32a20079eb1d57bd3c153888f8fd18c61f25204358a7d8aa10f71838b8f24890

Observation b2752897-a48e-4fcd-8676-bd76aaf0634e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.079176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.079176Z digest=sha256:4fe5fa9ce6e2fde60d271415cf6fdd3430ee52332d3f4eeacab19b5bf8602cee

Observation 802c7432-bfa4-4aa4-bf4b-795b6f5085a5 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.083942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.083942Z digest=sha256:6e1fe4d9e96696031f86299ec428bb7d7ab04d377eccc5cc310b8e18722c8695

Observation 21b18e6a-5855-4e9d-8e84-92dcd08b8b93 · outbound

This paper cites The History and Risks of Reinforcement Learning and Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The History and Risks of Reinforcement Learning and Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.089208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.089208Z digest=sha256:05593a965066b52fa9c491c154235830be6f9ebaf93390b39286f471abc4563c

Observation b79f6e8f-9af0-405d-870f-bf3606118fb6 · outbound

This paper cites The past, present and better future of feedback learning in large language models for subjective human preferences and values.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The past, present and better future of feedback learning in large language models for subjective human preferences and values

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.377687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.094552Z digest=sha256:2a6d757320b000a6b7091b7a3b031d8d14600a02aba810e95960f1240283e790

Observation 388ff225-b08c-4b3a-8bcf-c0d955670635 · outbound

This paper cites Artificial Intelligence, Values and Alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Artificial Intelligence, Values and Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.099490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.099490Z digest=sha256:f33f2d7f2c77c57dae3aa8aa0eeadd7121ed6ccbc0cd82c015bee9b55e5d3266

Observation 028d699f-485d-4526-867a-7036adcfa59f · outbound

This paper cites Artificial Intelligence, Humanistic Ethics.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Artificial Intelligence, Humanistic Ethics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.359427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.105715Z digest=sha256:55e836bdfd8e5a8a21074fe3e9125d4a668fd76f72e01ac2809c19e945a98f26

Observation 721b89fe-31a3-4c83-a215-fb417b6944ce · outbound

This paper cites Beyond Preferences in AI Alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Beyond Preferences in AI Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.116195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.116195Z digest=sha256:fe1cda4e6ea3f8ab3860175308120f960afbc633b6aba7783f79eb6d1898fc31

Observation 80e5b85b-db90-48d5-af76-5386352a2606 · outbound

This paper cites Meta Community Forum: Re- sults Analysis.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Meta Community Forum: Re- sults Analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.341743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.121237Z digest=sha256:c2b2529d48e2d6192971d48a6eff73f19f3f0a75f6cf655511dfc8c07129aa95

Observation e15e84d2-04ba-4681-b23e-f62263bc02dc · outbound

This paper cites STELA: a community-centred approach to norm elicitation for AI alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization STELA: a community-centred approach to norm elicitation for AI alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.125962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.125962Z digest=sha256:a5df9e57600d42be1abe0e2f2a25165e8fd2abcae3e2dd3993f64c26986979b6

Observation 59936292-d71c-4b0e-b898-45c3c8c5b2e2 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Constitutional AI: Harmlessness from AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.130971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.130971Z digest=sha256:adfbe6041b8dbecd2a608f15ec31876477dc15b5e5d329c372a28f875586db04

Observation db993abc-2c68-4cd6-8ce3-5b35f2fc6b72 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Improving alignment of dialogue agents via targeted human judgements

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.136182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.136182Z digest=sha256:f19b1a87c99a3218b88f2115aff3ab5c18435a95f286389924bb61ff861e07ba

Observation 99f768e8-7b04-482e-aa2b-8ad5b81e2964 · outbound

This paper cites Collective constitutional ai: Aligning a language model with public input.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Collective constitutional ai: Aligning a language model with public input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.141942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.141942Z digest=sha256:d1d80f427acdaf406647378465c623974986be5df323ef7c935341a3464f37d4

Observation 433669fe-1be5-4bf8-8e0f-5eb8d9ca4342 · outbound

This paper cites Introducing Meta Llama 3: The most capable openly available LLM to date, April.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Introducing Meta Llama 3: The most capable openly available LLM to date, April

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.311965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.147147Z digest=sha256:3342c1a7f6e706c67688761a4d91ddb8fd61b245b9986749ce3790b8bf121af3

Observation 14a7b8f0-b84a-4342-aa2d-8adb9e2c6fb8 · outbound

This paper cites MMToM-QA: Multimodal Theory of Mind Question Answering.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization MMToM-QA: Multimodal Theory of Mind Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.157581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.157581Z digest=sha256:a22b5f6f14e9177371cf36bb1ee8c9e34fc2c4d73ecda9a1d2308203dc13b960

Observation fdf7d294-b607-472a-bee8-ccc13d0ce288 · outbound

This paper cites ChatGPT’s weekly users have doubled in less than a year.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization ChatGPT’s weekly users have doubled in less than a year

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.277815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.162693Z digest=sha256:5ba57e3d8b388079fe9a858cb16a67e98af45ab2c8564de68d281585a44fec30

Observation 3bf9c8eb-cb4a-4d79-9527-d85eb53a5e0e · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization LaMDA: Language Models for Dialog Applications

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.167430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.167430Z digest=sha256:dc8d1e8fe8dc3d98f794926b4cacd61616d8be1c9e1bf5560fc1e8d34fffa7ab

Observation afd6c360-f8af-41b0-aee8-fa567431ebfd · outbound

This paper cites DICES Dataset: Diversity in Conversational AI Evaluation for Safety.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.172816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.172816Z digest=sha256:45d41ddd013f1f0ba92a692dda964de22b79c66260971037fe69c6a9a1876523

Observation 77e596d9-6cbc-4ec4-8aa7-bcae2448e9e5 · outbound

This paper cites Pretraining language models with human preferences.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Pretraining language models with human preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.177999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.177999Z digest=sha256:18e9f6c582a417bb4507d68ea21c266d28de2d4eee82fc56f8c9cd93189711f3

Observation a78eb4c6-df25-4f00-80ef-efc87adc010b · outbound

This paper cites alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization alignment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.249075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.182842Z digest=sha256:a876426864ffb6920454475bfa5ad6869e28da052213dc54d398b87ca2cd6f09

Observation 8686eea0-48ae-458c-86a8-cf66d1761691 · outbound

This paper cites Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.187916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.187916Z digest=sha256:1c435b844e1d6a7e5512acbfe381c2b2b92a3c7b7ee75ebc18c7f1ea13f0dc38

Observation 8e117d74-1c97-478c-af8f-2e7605ad7c3f · outbound

This paper cites The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.192853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.192853Z digest=sha256:072dae5001ce24b1f602de50b51bfc232d28ffaab41cca9f5bfc62c13d736b74

Observation a28e7541-daf6-4fef-8bc8-079c0e99f380 · outbound

This paper cites The Cultural Psychology of Large Language Models: Is ChatGPT a Holistic or Analytic Thinker?.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The Cultural Psychology of Large Language Models: Is ChatGPT a Holistic or Analytic Thinker?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.198392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.198392Z digest=sha256:f6a0820476cc72c0c70f28c378d700c5f8a1964b6e1bf2950c66d488281d9764

Observation da0a89d7-7cf0-4605-8415-302fdd4ef571 · outbound

This paper cites Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.203705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.203705Z digest=sha256:a6f9472c6ce9114033ed19461140917827bbe9442726bfb7ef71eff6c6289bd4

Observation 381a39c0-e688-4706-86aa-ce211ba081a8 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.209034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.209034Z digest=sha256:7c54712c62cd326967b97ddd412860157a4b4ffb824dad25aaa0107991bf9713

Observation f071d434-42d9-4bc0-85d9-6610397831de · outbound

This paper cites Personalized Language Modeling from Personalized Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Personalized Language Modeling from Personalized Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.214914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.214914Z digest=sha256:654ebc17c583d0c26ad56585e38358529e6678d4c5efa31c8b0bc4978e688826

Observation 1fed5acb-3a7c-47bd-853b-55d90103f283 · outbound

This paper cites Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.220449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.220449Z digest=sha256:55b7b04f956b11b48682041c923b55bfe866ae54a523a202d86935127e3c08fa

Observation 3b78df44-c63b-4647-8235-1d4c0525780a · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36, 2024.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.232281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.225923Z digest=sha256:1fcb770475e5537d2af0be35ffa6176fe38c427c783f081696a9917b6829b274

Observation ecf77a7b-1b82-40c2-9009-138632ef1698 · outbound

This paper cites Fine-tuning language models to find agreement among humans with diverse preferences.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Fine-tuning language models to find agreement among humans with diverse preferences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.230994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.230994Z digest=sha256:cf39d61d8631ad854ac37c176015b85c4f96166781be9888f2b3031a222b4be2

Observation 2d01beb3-7a78-4607-acee-3dd57276c533 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.235932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.235932Z digest=sha256:213c6f2394259ec47ae8fe73a0fba73b2eef55bd2ed0bad40c5bc9ecd0a896e8

Observation 340b4c90-159c-45f6-87e5-8086a3184a9c · outbound

This paper cites Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.241253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.241253Z digest=sha256:c4edb94ca1a1194c49d134c72c0fdea09e0f33eb6a14007bf453655b339c3ae8

Observation 23783648-074d-49db-887b-bd1fdb950640 · outbound

This paper cites Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.247150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.247150Z digest=sha256:04fcde6b251e8f8dab82fc17055a7c88b1fe8b7928c239d3a742b8647cb5f6f7

Observation 08fcd4e1-cc17-4c7d-a100-2797d16e8a90 · outbound

This paper cites Aligning Crowd Feedback via Distributional Preference Reward Modeling.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Aligning Crowd Feedback via Distributional Preference Reward Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.252782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.252782Z digest=sha256:17775af5b1b5c23b8e751dd65a761441105c499d3a284905d8b7fb9c94eaad32

Observation 76c581c7-7fd2-4bd6-9bfb-61de741da16d · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.258372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.258372Z digest=sha256:92fa808266f46f27a2f5027641a9bc90a00353d8460ccc190bdd19bc6dbde21c

Observation dc4ed378-108a-4b3d-b1da-7c481dec6011 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization WebGPT: Browser-assisted question-answering with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.264538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.264538Z digest=sha256:b89acc2a61304560625f36cc556289d8ccaba3cd772a5c9aa20539bc944ca691

Observation 27d43019-d72c-437d-bd35-d10a9dc4b073 · outbound

This paper cites synthetic-instruct-gptj-pairwise (revision cc92d8d),.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization synthetic-instruct-gptj-pairwise (revision cc92d8d),

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.203555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.271386Z digest=sha256:3c604f9fcc795108853f6b7572e117a180ed70a5581801bb7fdbf03c3bc155b4

Observation 2ade0842-68ab-4012-b155-8a9a76cbb5d4 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.285345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.285345Z digest=sha256:e2adede38384c9dc1fa7194b0346d2dfd807c92b2d7a6e76f3f8a5f652716be5

Observation 06542e24-3659-4559-ab7f-0bd9add0c4b6 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The Llama 3 Herd of Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.291078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.291078Z digest=sha256:bea55be52c09df9b2d5224d380f198c57c0831facf459be2074203f32910107a

Observation ac75e5fa-82e2-4ce3-ae80-ffa97cef70bc · outbound

This paper cites Transformers: State- of-the-art natural language processing.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Transformers: State- of-the-art natural language processing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.156056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.296646Z digest=sha256:c00cf27856e5373bd0d2378ad7541ba8faceb40a4395bda94d839524987084d7

Observation 396228b2-8607-415e-a327-5bec2dd767af · outbound

This paper cites Holistic Evaluation of Language Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.303089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.303089Z digest=sha256:6c6dfd10be285891a992b0df4532f0fd91752317765ad28323b5f2c2775a8ba8

Observation d7bab6bc-55b5-4a19-87c5-021f937b027b · outbound

This paper cites Argyle, Ethan C.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Argyle, Ethan C

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.137895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.308510Z digest=sha256:e560b7fc2ed1ee26bf5f35e3a301a1611e46a6e3dc8e371ebc81ebee8ba00522

Observation 06bdd5c2-9118-4f2d-9ee2-a7dbd851d771 · outbound

This paper cites Are large language models good annotators? In Proceedings on, pages 38–48.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Are large language models good annotators? In Proceedings on, pages 38–48

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.120199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.322035Z digest=sha256:3a42ed6295511efbb56c71c259488f64ff75a6e3d229ab606482ff8bc754348a

Observation ff1e3223-68a9-40a2-bd97-2f351a502601 · outbound

This paper cites Automated social science: Language models as scientist and subjects.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Automated social science: Language models as scientist and subjects

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.327536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.327536Z digest=sha256:af56ebaacd60733146416b1f9f81cc30c77ff15078eae5486a106afc041b16ba

Observation 7bf040e1-683c-4a15-8c3e-9923cb3bf3b1 · outbound

This paper cites Large language models that replace human participants can harmfully misportray and flatten identity groups.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Large language models that replace human participants can harmfully misportray and flatten identity groups

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.333381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.333381Z digest=sha256:2ae285b71b619fbcc50f12f38d75d1354b9fc0a72151a6ba60646b577332908d

Observation a1068a39-73d5-4815-8761-172725cef2d9 · outbound

This paper cites Out of One, Many: Using Language Models to Simulate Human Samples.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Out of One, Many: Using Language Models to Simulate Human Samples

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.315201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.315201Z digest=sha256:40c195f5f8796386dfeb0d638e4f5fc2a72367e929ab3e767441f149468dd3be

Observation 6da6ab26-ac28-4126-945e-024afc033184 · outbound

This paper cites Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.077496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.344513Z digest=sha256:9dde3a5464e8a70eaa525a03aaeded4061f920018027d886505be8a7c1ff0945

Observation 39f114fe-2bd6-4ba9-9612-161e00d9ce6a · outbound

This paper cites A Roadmap to Pluralistic Alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization A Roadmap to Pluralistic Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.356046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.356046Z digest=sha256:a86f768432b732b03fc3d6b6201005d9cf0cdbe9468e3fe8fab7ce8438b0d54d

Observation c3760899-b7e0-433e-b121-187ef371e613 · outbound

This paper cites Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971–30004.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971–30004

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.339375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.339375Z digest=sha256:6cab823bb0e993e2dd3785c11751ce8bea07c53ef08c536196ee59e41f1f1cc0

Observation 91f673de-6b3d-4e78-b56a-ffc829469bca · outbound

This paper cites The illusion of artificial inclusion.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The illusion of artificial inclusion

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T22:06:51.462091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.350182Z digest=sha256:6ce15067930de0511cc5ac6cf2a52cc24775baa434a5a58fe8aa5f6b1634bc95

Observation 8d6231ee-7d13-4e0b-b006-28e9bb0d1b04 · outbound

This paper cites doi: 10.1162/daed_a_01912.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization doi: 10.1162/daed_a_01912

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.110809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.110809Z digest=sha256:70b56c5c510c20d95ffc4bf44da70b667b39e24e595f7cd7aa5cf793f7513fa1

Observation 6de0a1df-d639-40e3-bb12-c5a52b5baf6e · outbound

This paper cites an unresolved cited work.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:06:52.185515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.277039Z digest=sha256:6c289808fb916afae238a15708c9b4942fd89bada241b37049d845674c512f52

Observation 14ef19e5-51d4-4da7-abc4-c3fa5effd67b · outbound

This paper cites an unresolved cited work.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:06:52.294454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T22:06:51.152203Z digest=sha256:e99e2cdf0b092b294fe1d24b6fec50f79128df295cf8108aa42d008a387fc9d1

Pith citing papers

Observation 7bad1392-2f41-4fc1-94b2-27e6f266a4c2 · inbound

What Do People Actually Want From AI? Mapping Preference Plurality cites this paper.

What Do People Actually Want From AI? Mapping Preference Plurality Beyond the Binary: Capturing Diverse Preferences With Reward Regularization

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:41:29.823462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T01:32:04.660400Z digest=sha256:880e04012f895667080ebbf2815b20950c1ee6292b00fcfd693710d5bad137e4