Pith. sign in

Paper Citation Record · LEDGER

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2509.03672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03672 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:54:13.607657Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:49:14.243931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:07:30.215801Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe8cde69-a40f-4ab0-9ee2-9c163d20786b · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences A General Language Assistant as a Laboratory for Alignment

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.716053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.716053Z digest=sha256:8da95a52bda4aeb74d1cc6e20d52aec0dc017c8a50ef8ffd4ec789636f50678d

Observation de8ce66d-7c79-458c-a04b-ec831e7f9bdb · outbound

This paper cites A Sharp Fannes-type Inequality for the von Neumann Entropy.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences A Sharp Fannes-type Inequality for the von Neumann Entropy

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:54:14.125679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:10.769694Z digest=sha256:c9eb681fcc23b96ac6918c688b76c6592cea9d643a392433cda058390d86e0d7

Observation 98653cfa-3c35-42ff-94b9-eda507bd7ace · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.832231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.832231Z digest=sha256:032c5d7772c51c50a8888aa3f64cf29b68fb060d1d1b20e07bf56642a815d093

Observation e687d5f9-1a48-4b56-acfc-308db9130456 · outbound

This paper cites Rank analysis of incomplete block designs: I.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.891723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.891723Z digest=sha256:ee205c38552a5cebe40274c5b5c8b520e0ef313613f99ef12e4cf9fead931481

Observation 97bb6238-1d2c-4042-b3e8-2ea3e4dddebe · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:10.974544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:10.974544Z digest=sha256:c17aa655bb3b8770ab83ca690363d031008440015dcd1324d6fdd83ee028b0a0

Observation d5241516-3b74-46d1-9351-296611749a36 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.055317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.055317Z digest=sha256:da2177b4fdca0e3c2f0e8929dabd7f1f7561b330696d7ce056e962dd1c439a4e

Observation 1fb46bd4-3130-4a4c-8d0c-9357b37ad186 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.144244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.144244Z digest=sha256:c8f9b5207b82711bb7604dd58c1f190d5aaeaa3c226cf1764b2e4af7bb2ee3aa

Observation f4181020-b7af-400b-acc7-19cf002d6b28 · outbound

This paper cites PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.230380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.230380Z digest=sha256:e4f1a350b490d4e5bc838b053a1ce47aae23006b75abb70364c6b2688a9a3b2d

Observation 60305593-17c6-4416-83df-aa8d0c54f42d · outbound

This paper cites The alignment problem: How can machines learn human values? Atlantic Books, 2021.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences The alignment problem: How can machines learn human values? Atlantic Books, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.339079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:11.305666Z digest=sha256:b4d10bbeda6c57defc5cee61c790c12d5dc8e260cb1683dfeae10e730aa2c6b8

Observation 936afcaa-c91e-4cd2-89f3-fb19b297db57 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.358214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.358214Z digest=sha256:afabfb0d1683c2ab486d9f81d3bddef333c5c0bf79f1ae3865f96bffda98d351

Observation 41632b7a-2b4f-4813-bf37-6d0ee4a79b09 · outbound

This paper cites Active preference optimization for sample efficient rlhf.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Active preference optimization for sample efficient rlhf

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.424298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.424298Z digest=sha256:dddd4d7a84a228c9d43fc1d6e350b90f0d29f887442f73377a83821798ab4348

Observation 88ca93d6-092d-48ca-af65-522c391164de · outbound

This paper cites When personalization meets reality: A multi-faceted analysis of personalized preference learning.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences When personalization meets reality: A multi-faceted analysis of personalized preference learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.466795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.466795Z digest=sha256:012c5a640a785332c125e928620e83c2172b023a5566ad2427d45662460367a3

Observation 1accefff-7cf8-4e75-a659-1d2262aba6fc · outbound

This paper cites Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.555391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.555391Z digest=sha256:79ced30c770c1227d133ba2464cb55972e14e8a3bb5317cfcc3c60c05e4b8f25

Observation 678c3fde-147b-4386-a1d5-01868278dfe5 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.639883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.639883Z digest=sha256:79e9739dd0dc5e6e21f057d461ed2f9800b8fabf08cad037eb16a4fb53e2180e

Observation 4b01892f-6150-4dd1-aecd-dc5de55d79c3 · outbound

This paper cites Provably feedback-efficient reinforcement learning via active reward learning.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Provably feedback-efficient reinforcement learning via active reward learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.314729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:11.724453Z digest=sha256:460bd9df06a1c38d75894f3bfbe31861c220d8e4e740d8e1372eb7e048a3ec5f

Observation 9e840360-9f7b-424c-835f-e8e8a47a2a0b · outbound

This paper cites Bandit algorithms.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Bandit algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.778287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.778287Z digest=sha256:60a05b36960da8f6b7ca1c59c1f67f02b1e39fc124a48338ce9e294e0c1e6edd

Observation 68bd6346-be17-4508-aee6-fc743a492faa · outbound

This paper cites Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.835948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.835948Z digest=sha256:0e083d7dceb31b44dc9f4ec3230ec2195e5439c2a3a9cc1cc5b8f4f9fe59fd89

Observation ddb95464-f354-4e29-985d-ce7d3544a8a2 · outbound

This paper cites Learning word vectors for sentiment analysis.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Learning word vectors for sentiment analysis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.288750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:11.925172Z digest=sha256:90e6b63bad46bbef9bdd71f30a8efab8346d66c82c17773a83b0d8b006e6ba53

Observation 6de299e3-5109-4788-9121-e654899f0bed · outbound

This paper cites Training language models to follow instructions with human feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Training language models to follow instructions with human feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:11.989754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:11.989754Z digest=sha256:a9233cccbef70b3dc994c683d79a27efa644dca988682c10bb868a1634cfb82c

Observation 20c57c66-534a-42b9-a12b-fb52ed52ceae · outbound

This paper cites Iterative reasoning preference optimization.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Iterative reasoning preference optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.079907Z digest=sha256:65652fc378afa15c2177605c54c8453a0fa1edf6f86c9ae76a77e32fe47fbf56

Observation 170584b6-cd83-4933-b97a-d3ddcbacea7a · outbound

This paper cites Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.137540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.137540Z digest=sha256:ceb6d1d72884ff06def43dbde4d871bc6f20fd8c767a03967913bddeab2e1ab1

Observation 0d3e55d7-8115-4545-83da-7e1eaffe14c7 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Direct preference optimization: Your language model is secretly a reward model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.237589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.237589Z digest=sha256:9cdc90dde561e93c5c06873f77f692f033fd54d57cf98462eba94900fb51c438

Observation ca89941e-85a1-443f-afb1-22990b46f596 · outbound

This paper cites Group robust preference optimization in reward-free rlhf.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Group robust preference optimization in reward-free rlhf

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.235770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.320047Z digest=sha256:553e29f9a43f271d3b30cf584cb01c634f84abed95d82257ded2823a9b88316b

Observation 3b9ce4ec-c2af-4127-ae71-2169cf9c2400 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Dueling rl: Reinforcement learning with trajectory preferences

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.219234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.412731Z digest=sha256:58505fa3e5c7a99dc1db237558a2fd357f0e6fbe4c1f0889c61643a51b5288d7

Observation ef35e711-f3b8-44ef-951b-25f36f72e3b4 · outbound

This paper cites Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.204135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:12.495223Z digest=sha256:8e605de18fe99fd47da0b258a99a4ad6143670d74f8434905297b6701258bf68

Observation fa21e778-d9e4-4fdd-9cc0-0aafdad75656 · outbound

This paper cites Proximal Policy Optimization Algorithms.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.534185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.534185Z digest=sha256:fe63bb3becbfd23df31c17f5f1fdc8790a09e25bc994ad727d09cefd9879628a

Observation 452cded4-fff2-4aa3-98aa-e5c967bb6dc6 · outbound

This paper cites Collective Choice and Social Welfare: An Expanded Edition.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Collective Choice and Social Welfare: An Expanded Edition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.626549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.626549Z digest=sha256:e05f8c7336d8c954b721323a2717fc53d87984366f4feb268df52fdf9b4d558d

Observation ab95073d-b7e5-4df8-95de-abd0a5a34da5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.794995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.794995Z digest=sha256:cd5d6cee2c3ddc6a473ae67cb959b75b25fecd697cf0032456830ffe246e1a35

Observation 7d07935c-e1dc-40c4-8720-48d0455f98e3 · outbound

This paper cites Robust multi-objective controlled decoding of large language models.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Robust multi-objective controlled decoding of large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.873574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.873574Z digest=sha256:084bcb2f60ee22ee34ad36d61607f748e79b98266c9909880703c14a10a62b40

Observation 1e85c555-f779-4e4c-bcd1-16de7045dfa7 · outbound

This paper cites Learning to summarize from human feedback.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Learning to summarize from human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:12.976902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:12.976902Z digest=sha256:347b8164b8cc513ac5e9e9534d58a046cbbde1dc576919d6b3038995464b091f

Observation 57da1d23-91bb-4077-aafa-9c65c0328894 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Aligning Large Language Models with Human: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.130765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.130765Z digest=sha256:4250eb5dde1ddb95c037204029324c9a2e869f68c13bf70924adfd87be8fdbb2

Observation 456bd312-3c97-43e2-bcc5-7cff9de0513a · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.249178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.249178Z digest=sha256:3684e534db465299efcc476476588ce0eb95d6daa09cacedbf86cf7908789607

Observation 50b4bcdc-9e2e-4605-a69f-c186ffb3f573 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL -constraint

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.188631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:13.373392Z digest=sha256:a441d34a30a0178cb945add35db014fc3ed94fbfe9c694ad0e3fbf49d2134fd9

Observation 96b151ef-cf89-42e9-823b-18e4492e0c01 · outbound

This paper cites Impact of representation learning in linear bandits.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Impact of representation learning in linear bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.173019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:13.481709Z digest=sha256:183090360d260196f8119f672be3caabdff373dad33372be2de945037e934657

Observation 243ea628-e798-41c4-acb3-3fda90368e0b · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.592468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.592468Z digest=sha256:ccaa3707235c02b6702f43a6d7d948dc735eabc011874a3e4532c01c9134b145

Observation 77f38854-3e60-4eee-9a0c-ac68f7e0abe5 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:54:14.158270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:54:13.607657Z digest=sha256:1e894c72e684dae324d8bf520af262db115a4dbc05e1c2a5313de9df46aa62b5

Pith citing papers

Observation d01e6178-4e4e-4035-9ad6-0c8a4e3310de · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.217449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:6f6e43f56144226629e610de56bb745bd1086c93476eb68984a29d7929503d06