Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.452310Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2501.00911.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.452310Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ec3479bc-16d1-42ff-a4ff-c498c8b31ee3 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b844550-bd51-45df-8c11-613835eda028 · outbound
Aligning LLMs with Domain Invariant Reward Models https://claude.ai/ Claude
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a3f310d-bb93-4017-8622-985e59cf9656 · outbound
Aligning LLMs with Domain Invariant Reward Models Wasserstein GAN
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a944a4-4d6f-4d3f-8f29-81a8644bf16b · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0b8ffb-f07f-4273-8905-0b6072b4c0e3 · outbound
Aligning LLMs with Domain Invariant Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd5e676-b502-475b-96b3-aa5f68fe30fd · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b63d2dc6-dc34-44d1-b8f6-9dc536e35640 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f22e4843-f4f6-4675-a54d-bd31bcd59cb4 · outbound
Aligning LLMs with Domain Invariant Reward Models The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e8f850-9bfc-4a31-b938-e890656a2175 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3beec7c8-733f-454b-8434-ebf1fdfb028b · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80268a2a-c2ba-4c37-bdda-e06be7ef805e · outbound
Aligning LLMs with Domain Invariant Reward Models Ustinova, Hana Ajakan, Pascal Germain, H
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0eff83a4-e1f8-4ef7-ba63-5a00b1299d15 · outbound
Aligning LLMs with Domain Invariant Reward Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8c36a4-6021-4a39-bfe8-28126657e856 · outbound
Aligning LLMs with Domain Invariant Reward Models Gemma: Open Models Based on Gemini Research and Technology
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d03db23-832b-4021-977d-f7bf771db709 · outbound
Aligning LLMs with Domain Invariant Reward Models Improved Training of Wasserstein GANs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525eec75-07ad-4631-9a9f-da21d0e2f561 · outbound
Aligning LLMs with Domain Invariant Reward Models Direct Language Model Alignment from Online AI Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9715775d-83cb-40a2-8121-0d37b211048c · outbound
Aligning LLMs with Domain Invariant Reward Models The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0482d9d0-e854-4409-83b1-de60c72f6608 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b571a44-facb-4844-b079-229b3f4fc963 · outbound
Aligning LLMs with Domain Invariant Reward Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2111e6-485b-465c-b7c6-982e37b8167e · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f719bc80-b061-4a16-9be3-dd5ca1ac8290 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5aa80a67-2fb3-40b9-af20-44ae63ebf13b · outbound
Aligning LLMs with Domain Invariant Reward Models UDALM: Unsupervised Domain Adaptation through Language Modeling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f121deb6-27f1-4e45-a3cb-39abe0b41336 · outbound
Aligning LLMs with Domain Invariant Reward Models Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399c1371-c5ea-41d9-900a-eb1a7ec50e34 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cac133d6-8dd6-4c6a-8397-d1f010ce696d · outbound
Aligning LLMs with Domain Invariant Reward Models Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67971a3-3e53-47bb-bb5f-339ef22d38be · outbound
Aligning LLMs with Domain Invariant Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2855cfac-44af-460d-9951-9d95e4373c8c · outbound
Aligning LLMs with Domain Invariant Reward Models Scalable agent alignment via reward modeling: a research direction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad49babe-b993-4957-9fee-352854386105 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation adb34e46-df8e-40f3-9fda-79761d91fc62 · outbound
Aligning LLMs with Domain Invariant Reward Models Preference Tuning For Toxicity Mitigation Generalizes Across Languages
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d1d84f-8ece-4bb0-8598-380f9625dd49 · outbound
Aligning LLMs with Domain Invariant Reward Models Decoupled Weight Decay Regularization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df829618-1316-49a3-ba2b-111a0c5aa761 · outbound
Aligning LLMs with Domain Invariant Reward Models No Language Left Behind: Scaling Human-Centered Machine Translation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9293c1f-65b5-46e0-81f3-3e1e4d1441dd · outbound
Aligning LLMs with Domain Invariant Reward Models https://chatgpt.com/ Chatgpt
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aad28e62-e728-4eea-bb76-3be7c1245df8 · outbound
Aligning LLMs with Domain Invariant Reward Models Training language models to follow instructions with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb05ed5-6849-4bc4-ad4b-7687b0f1e684 · outbound
Aligning LLMs with Domain Invariant Reward Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c82572-4120-4f94-971a-a25ff1d87f10 · outbound
Aligning LLMs with Domain Invariant Reward Models Proximal Policy Optimization Algorithms
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68407d2-bd69-4615-b5b9-025746aa746f · outbound
Aligning LLMs with Domain Invariant Reward Models Aligning Language Models with Demonstrated Feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8d7799-1e80-43f4-8d2a-29b83943b9e6 · outbound
Aligning LLMs with Domain Invariant Reward Models Wasserstein Distance Guided Representation Learning for Domain Adaptation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d92ae9d-cbe4-41d3-96c7-d9cd1c25a657 · outbound
Aligning LLMs with Domain Invariant Reward Models Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de6c780-9bdd-4dee-835c-7e2d8618b558 · outbound
Aligning LLMs with Domain Invariant Reward Models Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef303f7c-fb2a-4fb1-8fde-c11ed3a7cf1d · outbound
Aligning LLMs with Domain Invariant Reward Models Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15eaf064-48e0-4a6e-8d5e-c6c2f796e857 · outbound
Aligning LLMs with Domain Invariant Reward Models Chernova, and Dhruv Batra
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d9e7e8de-fa6f-4e8d-a26d-4cccde6dc901 · outbound
Aligning LLMs with Domain Invariant Reward Models Deep Domain Confusion: Maximizing for Domain Invariance
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e561a12-24a5-4b02-b30d-94e811486f54 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adc45c0-0464-4e51-a507-fb1f30304374 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a05d9238-5b9f-42d9-b0e6-ff3ebf2d2d76 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1340b498-31a0-45e7-b283-1ec86018401b · outbound
Aligning LLMs with Domain Invariant Reward Models Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97b509e2-18b3-48b6-9e74-8dfa8fc741ee · outbound
Aligning LLMs with Domain Invariant Reward Models CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99232992-40d2-465d-94e1-27471d260e93 · outbound
Aligning LLMs with Domain Invariant Reward Models Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e724335a-a9ae-42b5-9421-b8240cf682a8 · outbound
Aligning LLMs with Domain Invariant Reward Models Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0496f569-b2bf-4586-9b6e-075691c6365c · outbound
Aligning LLMs with Domain Invariant Reward Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a406a80-2238-4f13-bd0a-8ba6f59e3246 · outbound
Aligning LLMs with Domain Invariant Reward Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f39061c4-0a81-46a2-bb04-17b0c58d1b40 · outbound
Aligning LLMs with Domain Invariant Reward Models Fine-Tuning Language Models from Human Preferences
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b527aeef-d4b6-463d-9df6-3949ecd8a411 · outbound
Aligning LLMs with Domain Invariant Reward Models online" 'onlinestring :=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1adb606a-3623-4c4a-a5c1-dfb6f422aee9 · outbound
Aligning LLMs with Domain Invariant Reward Models write newline
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.