Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:11.905544Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2501.09254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:15:11.905544Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.207057Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:49:13.182463Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a0ec520-0d0d-45c1-97ad-35ebf5e33ec5 · outbound
Clone-Robust AI Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4888a7-161e-4f2e-8048-a8ee1d2e01db · outbound
Clone-Robust AI Alignment Note that the function log er(x1) er(x1)+er(x2) is strictly convex in r(x1) and r(x2) as shown in Siththaranjan et al
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1579c556-f111-4291-9612-11d78047b2d0 · outbound
Clone-Robust AI Alignment MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2e9106-73c6-480a-a06b-f762c21d181f · outbound
Clone-Robust AI Alignment Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be4ed5f-9d33-49e9-84a2-d4acb4b8c731 · outbound
Clone-Robust AI Alignment Mapping Social Choice Theory to RLHF
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a273bc8b-778e-4e97-b164-b19dbe6c9da2 · outbound
Clone-Robust AI Alignment Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0b914654-272d-43a5-a17c-26fbff865f8f · outbound
Clone-Robust AI Alignment Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1addc87a-f906-4f9c-b46f-643a8576225f · outbound
Clone-Robust AI Alignment Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d3818d-8a1a-445f-bb9a-ecb28acf6327 · outbound
Clone-Robust AI Alignment Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d83fb8-77b6-408c-8023-f91cc88b43df · outbound
Clone-Robust AI Alignment A Roadmap to Pluralistic Alignment
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce193b1-1f6d-411d-8c9e-e20b03b610f8 · outbound
Clone-Robust AI Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a84e697-4de5-473b-8299-cc0681639635 · outbound
Clone-Robust AI Alignment On the identifiability of mixtures of ranking models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead4f0c5-aa75-4294-afdc-61c37b25fd27 · outbound
Clone-Robust AI Alignment Provable Multi-Party Reinforcement Learning with Diverse Human Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58808483-d97f-4e6c-9d95-f8da4d7ea6a8 · outbound
Clone-Robust AI Alignment Fine-Tuning Language Models from Human Preferences
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7ed927-ba3c-4554-b2b4-9bb1683bfeac · outbound
Clone-Robust AI Alignment Robust Reinforcement Learning from Corrupted Human Feedback
Reference 1952
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8821d90-dbfc-4d5e-bd97-b1fa53cf5ae3 · outbound
Reference 1987
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 274973ab-cbae-4819-b5e4-62b3e0f81993 · outbound
Clone-Robust AI Alignment Axioms for AI Alignment from Human Feedback
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15fee05c-a5c9-4b3a-ae2e-6c85836202ef · outbound
Clone-Robust AI Alignment Like us, Xu et al
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 766b1459-b304-4282-8234-ec931e14f67e · outbound
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 042204a1-9086-4a37-a54d-45d1a504af0c · outbound
Clone-Robust AI Alignment Corruption Robust Offline Reinforcement Learning with Human Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02e70172-0199-43fc-b90b-06d4001b91e1 · outbound
Clone-Robust AI Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30a5fa6-f7a2-4596-b4af-5a9f682e40df · inbound
Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Clone-Robust AI Alignment
Reference 149
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9158f3a-d4a9-4615-8655-ff963d98d703 · inbound
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Clone-Robust AI Alignment
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d89933d7-f66e-471e-b803-6fe0d74093e8 · inbound
Internal Pluralism and the Limits of Pairwise Comparisons Clone-Robust AI Alignment
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 716378cb-94ab-40ae-919d-bf78e599f517 · inbound
Internal Pluralism and the Limits of Pairwise Comparisons Clone-Robust AI Alignment
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.