Pith. sign in

Paper Citation Record · LEDGER

Aligning LLMs with Domain Invariant Reward Models

As of 11 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2501.00911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00911 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.452310Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec3479bc-16d1-42ff-a4ff-c498c8b31ee3 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.434617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.164501Z digest=sha256:9374513420fe252a9658941413977bd41e4422fd255605a61117de9fe8597166

Observation 4b844550-bd51-45df-8c11-613835eda028 · outbound

This paper cites https://claude.ai/ Claude.

Aligning LLMs with Domain Invariant Reward Models https://claude.ai/ Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.417746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.170691Z digest=sha256:069d0e785327842c4c4f37e5dfcb98c7fa44f97e370b716689248580739c2336

Observation 4a3f310d-bb93-4017-8622-985e59cf9656 · outbound

This paper cites Wasserstein GAN.

Aligning LLMs with Domain Invariant Reward Models Wasserstein GAN

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.176006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.176006Z digest=sha256:accbe9dfba3c998128d52afe8586bcc0e9c0b0536bb3cb96fcfedb7849bcfffa

Observation 22a944a4-4d6f-4d3f-8f29-81a8644bf16b · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.181779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.181779Z digest=sha256:20113543f5d0938ef2f56dd709f850bc095c32030cb3bf255bb9ca555dd1c6bb

Observation 4a0b8ffb-f07f-4273-8905-0b6072b4c0e3 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Aligning LLMs with Domain Invariant Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.187680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.187680Z digest=sha256:9f553e8de16ad029ea2ffada6216bac0255c3a72afd8a8bff19597c7fcdca4b1

Observation 9fd5e676-b502-475b-96b3-aa5f68fe30fd · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.390203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.198606Z digest=sha256:d7b17dc8ceac4c7dcc4bbe86cb325baccadb77bccfe7d3a57bb17bfd7c6735e5

Observation b63d2dc6-dc34-44d1-b8f6-9dc536e35640 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.370858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.204272Z digest=sha256:e0cd22c55147564270e630718cedb56c13d5eb78cfe9c44135241fae35eb5e8d

Observation f22e4843-f4f6-4675-a54d-bd31bcd59cb4 · outbound

This paper cites The Llama 3 Herd of Models.

Aligning LLMs with Domain Invariant Reward Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.209009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.209009Z digest=sha256:29b0fd1d3076f78a984d1d9f66dc102dd7e8a2fe30ab261e0a81c42491987fdd

Observation 51e8f850-9bfc-4a31-b938-e890656a2175 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.215495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.215495Z digest=sha256:4eb00cb08604adaf447ae22427ea9f77c870656a997378abeb404e49e6fadbcc

Observation 3beec7c8-733f-454b-8434-ebf1fdfb028b · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.339535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.220549Z digest=sha256:d86aba2e7b15a14436e7baac3f61fb66334d384752281471563aacc151036ecb

Observation 80268a2a-c2ba-4c37-bdda-e06be7ef805e · outbound

This paper cites Ustinova, Hana Ajakan, Pascal Germain, H.

Aligning LLMs with Domain Invariant Reward Models Ustinova, Hana Ajakan, Pascal Germain, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.321934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.225250Z digest=sha256:1e30c3121f2367f695a4e3db0a7ffb218a1c5e5b1134764fa6257ee129e2303f

Observation 0eff83a4-e1f8-4ef7-ba63-5a00b1299d15 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Aligning LLMs with Domain Invariant Reward Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.230638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.230638Z digest=sha256:c64d9d050e88eb8a8c352d32fddb819f504e978efe18f77027b8b5d7e0e36e47

Observation 2d8c36a4-6021-4a39-bfe8-28126657e856 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Aligning LLMs with Domain Invariant Reward Models Gemma: Open Models Based on Gemini Research and Technology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.235972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.235972Z digest=sha256:56779a3d617d6336baacb0184d03fcacc9f23456b99444299d027fb3bf9b27fa

Observation 7d03db23-832b-4021-977d-f7bf771db709 · outbound

This paper cites Improved Training of Wasserstein GANs.

Aligning LLMs with Domain Invariant Reward Models Improved Training of Wasserstein GANs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.241129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.241129Z digest=sha256:a52b8bf4d1648a11d8b90b23f4852f1bdbfed4602a03ff1f75cc02509d78f6b3

Observation 525eec75-07ad-4631-9a9f-da21d0e2f561 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Aligning LLMs with Domain Invariant Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.246781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.246781Z digest=sha256:243a83ccaae9f1b755fd580ea9b735eac96ffa8c8246aff7bee6612edd69eb03

Observation 9715775d-83cb-40a2-8121-0d37b211048c · outbound

This paper cites The Unreasonable Effectiveness of Easy Training Data for Hard Tasks.

Aligning LLMs with Domain Invariant Reward Models The Unreasonable Effectiveness of Easy Training Data for Hard Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.252155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.252155Z digest=sha256:96c8dbafd179f76b1f3482826d726701d3caec8b0701f9ce1bd114ded16664aa

Observation 0482d9d0-e854-4409-83b1-de60c72f6608 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.302982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.257783Z digest=sha256:6765929951a5f0adb402c698d62234e5c656fe2bf86269f2ce6ac678c99d1ce2

Observation 1b571a44-facb-4844-b079-229b3f4fc963 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligning LLMs with Domain Invariant Reward Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.262601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.262601Z digest=sha256:ffb9b3da84356d78e5e3c7397627cc6b21c19540cd6a4e1ba7b3e9c0d377e8f4

Observation 1c2111e6-485b-465c-b7c6-982e37b8167e · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.285434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.267497Z digest=sha256:29b9b6b67310b994cea26fa079b0b940e11d33bb9b9cf10b81d433c96e733523

Observation f719bc80-b061-4a16-9be3-dd5ca1ac8290 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.268035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.272533Z digest=sha256:a012f5a6212c36694fb8699179ac84c1d2b0a08937ab9dae43cb70cb738f0c96

Observation 5aa80a67-2fb3-40b9-af20-44ae63ebf13b · outbound

This paper cites UDALM: Unsupervised Domain Adaptation through Language Modeling.

Aligning LLMs with Domain Invariant Reward Models UDALM: Unsupervised Domain Adaptation through Language Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:45:12.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.277915Z digest=sha256:e8730f7eeb9b64a4d368f326126e72d88bdebee82d79ff057ddc3401bc41f210

Observation f121deb6-27f1-4e45-a3cb-39abe0b41336 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

Aligning LLMs with Domain Invariant Reward Models Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.283131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.283131Z digest=sha256:32519eba0b07d9119cea4195751c9dd05e917311343d3bb52f8534efb0f92b55

Observation 399c1371-c5ea-41d9-900a-eb1a7ec50e34 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.247219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.288497Z digest=sha256:bc6fdb9ce60e8ec60f5fd584214b2568d77b38047f0536b43f06badf88ee7db5

Observation cac133d6-8dd6-4c6a-8397-d1f010ce696d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Aligning LLMs with Domain Invariant Reward Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.293458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.293458Z digest=sha256:44f9f3c02593c2a03f392dd90dbce521cbd9d07d656c390514309885e4a2347b

Observation f67971a3-3e53-47bb-bb5f-339ef22d38be · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Aligning LLMs with Domain Invariant Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.299196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.299196Z digest=sha256:94a63b02fa08e48fb295bfde5d561876f9a8500ab9b809a8701019bd4d27d7f5

Observation 2855cfac-44af-460d-9951-9d95e4373c8c · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Aligning LLMs with Domain Invariant Reward Models Scalable agent alignment via reward modeling: a research direction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.304282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.304282Z digest=sha256:a1660821a544942217d6464197058426c0f643cb6761903b92ea95b3d690d891

Observation ad49babe-b993-4957-9fee-352854386105 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.228994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.309736Z digest=sha256:ac84b000e08b62f18eca4b788f3d6d45f52cae55bc360906dc69913b2aa2a8a5

Observation adb34e46-df8e-40f3-9fda-79761d91fc62 · outbound

This paper cites Preference Tuning For Toxicity Mitigation Generalizes Across Languages.

Aligning LLMs with Domain Invariant Reward Models Preference Tuning For Toxicity Mitigation Generalizes Across Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.315612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.315612Z digest=sha256:4e64a496f924174a9f703a10f7cb2fa5231168bbedc44e196424d95c9d9e97d5

Observation 73d1d84f-8ece-4bb0-8598-380f9625dd49 · outbound

This paper cites Decoupled Weight Decay Regularization.

Aligning LLMs with Domain Invariant Reward Models Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.322212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.322212Z digest=sha256:05838b07d4b3aa3cc20b54ce008fd6052ab6ac1106564da1e5690c513aeaa6ba

Observation df829618-1316-49a3-ba2b-111a0c5aa761 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Aligning LLMs with Domain Invariant Reward Models No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.326988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.326988Z digest=sha256:9d6e247617eba97394fe77192783404027b013aa20ec80795eb0d43b7e93b0af

Observation c9293c1f-65b5-46e0-81f3-3e1e4d1441dd · outbound

This paper cites https://chatgpt.com/ Chatgpt.

Aligning LLMs with Domain Invariant Reward Models https://chatgpt.com/ Chatgpt

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.211159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.332430Z digest=sha256:160a425ce9dece57182d540d7253534f92d299347dd5d4632af92e1afc396906

Observation aad28e62-e728-4eea-bb76-3be7c1245df8 · outbound

This paper cites Training language models to follow instructions with human feedback.

Aligning LLMs with Domain Invariant Reward Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.338619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.338619Z digest=sha256:1eb99ec36a3d23173435d5011951ccb66ad1d9003dff5e867b984701919020fd

Observation 8eb05ed5-6849-4bc4-ad4b-7687b0f1e684 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Aligning LLMs with Domain Invariant Reward Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.344286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.344286Z digest=sha256:8a7272c3670242d73669def080145b6601d7b94e0c421f49baf81b92ed4bb375

Observation b2c82572-4120-4f94-971a-a25ff1d87f10 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Aligning LLMs with Domain Invariant Reward Models Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.349688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.349688Z digest=sha256:3047298e5781774638a35eff4c32d6e0d8684fd91add49fcbea680888875df5e

Observation c68407d2-bd69-4615-b5b9-025746aa746f · outbound

This paper cites Aligning Language Models with Demonstrated Feedback.

Aligning LLMs with Domain Invariant Reward Models Aligning Language Models with Demonstrated Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.355146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.355146Z digest=sha256:6b59dbba866f91c43c906104ca537d791fa309663f03c57f308262bac4ca1091

Observation bb8d7799-1e80-43f4-8d2a-29b83943b9e6 · outbound

This paper cites Wasserstein Distance Guided Representation Learning for Domain Adaptation.

Aligning LLMs with Domain Invariant Reward Models Wasserstein Distance Guided Representation Learning for Domain Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.361165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.361165Z digest=sha256:de9accea466d815ac5ef2014e2e990c641ce7ff5609a3c67144ccd992cdf9ec6

Observation 9d92ae9d-cbe4-41d3-96c7-d9cd1c25a657 · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Aligning LLMs with Domain Invariant Reward Models Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.367288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.367288Z digest=sha256:42ec4ea7502529fe7df697313bde614bcc7451041f4485cbdfc8015b48652f67

Observation 2de6c780-9bdd-4dee-835c-7e2d8618b558 · outbound

This paper cites Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment.

Aligning LLMs with Domain Invariant Reward Models Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.372847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.372847Z digest=sha256:8b3d488fa29cd7c4fdee2a25cfc0f46ab78bbb52b42a50045a87d2cd9bac805b

Observation ef303f7c-fb2a-4fb1-8fde-c11ed3a7cf1d · outbound

This paper cites Causal Confusion and Reward Misidentification in Preference-Based Reward Learning.

Aligning LLMs with Domain Invariant Reward Models Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.378181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.378181Z digest=sha256:02640b502460b0fde4cf9e6f1fd23a85b6e8a7680ecfe3f49b03f6dbae709b6d

Observation 15eaf064-48e0-4a6e-8d5e-c6c2f796e857 · outbound

This paper cites Chernova, and Dhruv Batra.

Aligning LLMs with Domain Invariant Reward Models Chernova, and Dhruv Batra

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:45:13.181683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.383414Z digest=sha256:049b0d3853c67a036eb38616ba36a40991a9f7a06481302f53a2d74fe92f9c93

Observation d9e7e8de-fa6f-4e8d-a26d-4cccde6dc901 · outbound

This paper cites Deep Domain Confusion: Maximizing for Domain Invariance.

Aligning LLMs with Domain Invariant Reward Models Deep Domain Confusion: Maximizing for Domain Invariance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.388655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.388655Z digest=sha256:8f5073bc370dfb8ed799e5c086b78c9265579a017f6c196b9510fc0c0f993fa2

Observation 9e561a12-24a5-4b02-b30d-94e811486f54 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.393710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.393710Z digest=sha256:96605df90c043bc9efcef73a0dcf751a89197c48033a0ab3e467bd3804ea3cf9

Observation 6adc45c0-0464-4e51-a507-fb1f30304374 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.161418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.398638Z digest=sha256:c737d2b6d3f27dcc99c21d6f431913959ac594438ceafa4766b8f13af3f40316

Observation a05d9238-5b9f-42d9-b0e6-ff3ebf2d2d76 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.141030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.404542Z digest=sha256:5e559eeb24ccdd04e3f38b26f90dfc31eab7a9c6a80fae8568145d764f12f5a8

Observation 1340b498-31a0-45e7-b283-1ec86018401b · outbound

This paper cites Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment.

Aligning LLMs with Domain Invariant Reward Models Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:45:12.593278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.409531Z digest=sha256:aab75dfae0a5749c5f82cca30449ef0659aec9c31e8a85b45bf82ea0b85b1379

Observation 97b509e2-18b3-48b6-9e74-8dfa8fc741ee · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

Aligning LLMs with Domain Invariant Reward Models CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.414979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.414979Z digest=sha256:736061ef221ebaf0aef672de467b21ab3da0489ac9a556b621f894c28b32a24a

Observation 99232992-40d2-465d-94e1-27471d260e93 · outbound

This paper cites Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs.

Aligning LLMs with Domain Invariant Reward Models Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.420151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.420151Z digest=sha256:a896ff64da568827b5b774cafab0dd0378b1be8cb2bea3d44f9bd4133d700083

Observation e724335a-a9ae-42b5-9421-b8240cf682a8 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Aligning LLMs with Domain Invariant Reward Models Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.425370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.425370Z digest=sha256:cdbdc47ef8aad2bf8bd100a04d58d7e06b8639826ed9e34d07f5268d9db798ca

Observation 0496f569-b2bf-4586-9b6e-075691c6365c · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Aligning LLMs with Domain Invariant Reward Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.431407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.431407Z digest=sha256:c1edda893a9dd2c8b4857ba2f00afb0355236c32d989652ad579b32ed060c354

Observation 4a406a80-2238-4f13-bd0a-8ba6f59e3246 · outbound

This paper cites an unresolved cited work.

Aligning LLMs with Domain Invariant Reward Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:45:13.123721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:45:12.436817Z digest=sha256:ef0761ced91cc878e6a33b4a02f1c55feb45582ac79c35d1b0338d06e1fd9dcc

Observation f39061c4-0a81-46a2-bb04-17b0c58d1b40 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Aligning LLMs with Domain Invariant Reward Models Fine-Tuning Language Models from Human Preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.441631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.441631Z digest=sha256:c4779625ba45cf79d23528b5ece1247630755ac147e2efa8723cc3de647e0379

Observation b527aeef-d4b6-463d-9df6-3949ecd8a411 · outbound

This paper cites online" 'onlinestring :=.

Aligning LLMs with Domain Invariant Reward Models online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.446752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.446752Z digest=sha256:2aa1abfa44cba074cdc830a52cdbc9866b4dcb254f7fdec982a265020dd6bff3

Observation 1adb606a-3623-4c4a-a5c1-dfb6f422aee9 · outbound

This paper cites write newline.

Aligning LLMs with Domain Invariant Reward Models write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.452310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.452310Z digest=sha256:9608166be5495824c7928b69e7fbb41330a64965647c41e32e22d6e5e1619458

Pith citing papers

No inbound Pith citation observations are available.