Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:04.530071Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2506.07165.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:04.530071Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T14:02:11.084514Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T14:02:39.792894Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 509ef113-699f-40ce-8791-51ccddbc9a56 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a775cc4-d1d3-4e49-8a9b-cb25c3ccb4a9 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models - (2) Acknowledges both but slight deviations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1a3283-9453-4c75-b10b-7acce1c8d99c · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71380829-229d-44c8-9d61-8c582f171bdb · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Based the instruction following rule and given my answer to an instruction, your role is to provide specific and constructive score for me
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86162f79-067f-4f88-b0c0-d2f4e8348557 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4aab7911-1989-4d87-9051-75447098113c · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbff26fe-c0c9-46e4-bedc-349cecf533f1 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 603471d7-87ff-4375-9720-010016c55d82 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fbb7736-bb7b-42f0-86de-4dcedd084167 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Al ter na tively, you can use a nav iga tion app like Google Maps or Waze to get the most ac cu rate and up -to -date di rec tions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6c9871d-7ba5-4d27-ba1d-24fa4ab96c53 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6ca1b9-2929-4fd4-a204-d7da3e8cf2d7 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e92826-9fb6-46ae-8ebc-a6341bdc5fa9 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed371d59-8070-42de-9b40-2ce71ec1f177 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response completely missed the essence of what the user wanted
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cd52a9f-1b84-4ad1-a6bd-9cbb8fd30d50 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd53e2c5-9267-4541-9df4-ee050e8602fb · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response did not fully satisfy what the user was looking for
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa9b1903-cb42-4c0a-814c-254e76fae5c4 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbc3e889-c355-4e81-9ff2-9165a1ad4ad2 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Based the helpfulness rule and given my answer to an instruction, your role is to provide specific and constructive score for me
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32325681-7ec1-4dd6-9bdf-f0ea1f3c9990 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models All information provided is wrong, false or hallucinated
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52821ed3-340a-4e97-a005-215c70773b3c · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response may contain multiple instances of hallucinations, false information, misleading information, or irrelevant information
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58aa1c99-ebb1-497d-9d3e-ebc5d1b74f66 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response may miss some details, contain misleading information, or minor hallucinations, but is more or less aligned with what the prompt asks for
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7fd95f00-d9a4-45b7-9049-a5a45085108f · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models It contains no misleading information or hallucinations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 790f3e81-e2ec-499a-aab5-3cfcc4769e69 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models preference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0011564e-56b0-4aab-abea-a22f639a26d9 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f29bfbe-07b8-4c9f-a11d-1b34ce1daa2a · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dce37129-54b0-45fb-8324-9574dd80a10a · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4af259a0-678c-44b6-8a77-ec21b6b0bcc1 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48b3e66d-2b68-455e-8ba5-0529e032da40 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40e32249-f472-4723-8fff-fd3817ceeb14 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models “markdown“‘markdown. This is an example of a code block in Markdown. You can see that it is formatted to look like it’s not part of the regular text flow
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d84453a-c1c0-4a46-8b93-550b18215763 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76909b5b-1482-4ff9-8e44-9f0fcfd9bfd7 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebadc2cd-9ee1-40cb-b856-13ae76e0a176 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b862183f-ab10-4734-86ba-1069bdc073a7 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41fc8d2a-bdb7-4ad1-a839-ac4003e13cda · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46f191fa-7d8e-42ec-a79a-c6b79afd8bfe · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8905fa09-1251-4ee1-ad62-a8170e0d10ff · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0666dcdc-6eb4-4a53-9d2a-e916603e2a7f · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b3002d7-2969-41ed-bdf5-5b0cce6667fe · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76cc2232-739b-4d7b-8197-e6c39dab316a · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d97ec2f6-8e34-480a-9c64-b69b5dd026f1 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a13a5ca8-7485-485b-9904-bfad94f57cc4 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Reference 1027
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c276b0-1f1d-4f37-ba7b-f54921c2139b · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models KTO: Model Alignment as Prospect Theoretic Optimization
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46dcb65d-320c-4102-bf03-cf7d6de321d5 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b655c1f6-dca9-4a28-bdb0-b8ca25589367 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46af202c-d500-4828-b593-0f2cd9e422e8 · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bi- lal Piot, Rémi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91b171db-f386-42d6-88ab-6aa8242fb34e · outbound
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Anirudhan Badrinath, Prabhat Agarwal, and Jiajing Xu
Reference 4455
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f8dbf7d-2cd1-4c5d-8052-3748243927f9 · inbound
Failure Modes of Maximum Entropy RLHF AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.