Pith. sign in

Paper Citation Record · LEDGER

Optimal Transport for LLM Reward Modeling from Noisy Preference

As of 7 August 2026, this Paper Citation Record lists 100 of 249 outbound references and 0 inbound Pith citation observations for arXiv:2605.06036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06036 v1

Coverage vector

measured 100 of 249 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T14:10:49.634358Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 249 outbound references displayed

  • verified exact2
  • verified fuzzy56
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef82d082-58f7-46aa-b844-715c66e05766 · outbound

This paper cites Instance-dependent label-noise learning with manifold-regularized transition matrix estimation.

Optimal Transport for LLM Reward Modeling from Noisy Preference Instance-dependent label-noise learning with manifold-regularized transition matrix estimation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.752084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:c8a251ff802226d8555128a2f432c49d74fdc395dd21d30f9a3dabf5c64c125c

Observation 716b047e-b429-4f28-98f4-31adf64942d9 · outbound

This paper cites Class-dependent label-noise learning with cycle-consistency regularization.

Optimal Transport for LLM Reward Modeling from Noisy Preference Class-dependent label-noise learning with cycle-consistency regularization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.741408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:f3017ffacded7cfc560fbe307afd04beb19681637eb8211449d63c0710930f07

Observation d9940485-070f-4aa2-bed8-d87a54cd8fda · outbound

This paper cites Joint distribution optimal transportation for domain adaptation.

Optimal Transport for LLM Reward Modeling from Noisy Preference Joint distribution optimal transportation for domain adaptation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.760077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:7ca46fd34e5e1765fb8cc2f22ba7df59e170ef7215ad18cf46ae644cfbeac9da

Observation 85b27b5b-19ee-4a38-9ea4-7a0dc05c3aa1 · outbound

This paper cites Unbalanced minibatch optimal transport; applications to domain adaptation.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unbalanced minibatch optimal transport; applications to domain adaptation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.443247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4b919817e1199cb87a5cf57356bfea0ee12878206f7a15ddd2de4026eb18f560

Observation fa7553cc-59d8-43ab-8147-44c1e1aa482d · outbound

This paper cites A survey on llm-as-a-judge.

Optimal Transport for LLM Reward Modeling from Noisy Preference A survey on llm-as-a-judge

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.304146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:3faadd9fb32a3bcc6ed838de8bcc09bea93ef9402faf19801308c82daa904bf7

Observation c2d6b223-8d44-44e6-be33-6b0fbc8fd9de · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Optimal Transport for LLM Reward Modeling from Noisy Preference Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.363637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:f81b6bd60badb985575fc2847b49ce670094cf59b3b79ad60b993f781001ce32

Observation 0193d076-c9a8-4df1-a938-47b981dee321 · outbound

This paper cites Co-teaching: Robust training of deep neural networks with extremely noisy labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference Co-teaching: Robust training of deep neural networks with extremely noisy labels

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.300802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:8ba07fcc2afbcd37e74d4a59c3c8d50b32bd0c9f2f25834694124e012710c6bc

Observation 0db68485-3333-4ae0-8f9f-debff866412f · outbound

This paper cites Co-teaching: Robust training of deep neural networks with extremely noisy labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference Co-teaching: Robust training of deep neural networks with extremely noisy labels

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.561685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:0d99218635e13eb20d0f21f0784a66b2582b3a7b1d25a359100ecb58a7031a84

Observation 30c19581-d20c-474f-901e-13588a611e06 · outbound

This paper cites A survey on the role of crowds in combating online misinformation: Annotators, evaluators, and creators.

Optimal Transport for LLM Reward Modeling from Noisy Preference A survey on the role of crowds in combating online misinformation: Annotators, evaluators, and creators

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.347194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:554b07a90eaddfa32b185fbc4590c1322968e83f80f9d944758db53c725a8254

Observation 25492405-b697-45e2-b27a-2c3799a2250f · outbound

This paper cites A., Zhou, J., Wang, K., Li, B., et al.

Optimal Transport for LLM Reward Modeling from Noisy Preference A., Zhou, J., Wang, K., Li, B., et al

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.576992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:d2bc51ab9083781d85d10a0104a01ec078e41ab764e3cbfd3145be1f2b4c07f8

Observation 021cd081-7240-4594-9111-c850c04da875 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.287270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4d5870b749d409e6a26285b0c5d8634f71c2fd12b1c0d2c5d32984fed9b19bb6

Observation 932e336c-b758-4072-ac7f-e85e7e504dbc · outbound

This paper cites Instance-dependent label distribution estimation for learning with label noise.

Optimal Transport for LLM Reward Modeling from Noisy Preference Instance-dependent label distribution estimation for learning with label noise

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.241869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:94913574d12ffe7e03fb922f0e8bb46d4aea113df09fc89810e26f08f5f283ac

Observation 14365cfe-8f18-4d6b-92cb-2cd7cc8861cb · outbound

This paper cites Learning the latent causal structure for modeling label noise.

Optimal Transport for LLM Reward Modeling from Noisy Preference Learning the latent causal structure for modeling label noise

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.342816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4548b11d30f5f099327835be625a73bf99b9ba1b0c87b80024e1f265fe154cbd

Observation f96cd13e-6039-4355-90e4-705aa9cc11a5 · outbound

This paper cites Curvature-balanced feature manifold learning for long-tailed classification.

Optimal Transport for LLM Reward Modeling from Noisy Preference Curvature-balanced feature manifold learning for long-tailed classification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.398151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:70d9ff5f38e7fcd71721d3ff788a8bdf9c001bdd2409dacf1458d0f5e8ecb2af

Observation 70140c73-b50f-4cb7-ad90-98526e58a6b2 · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Optimal Transport for LLM Reward Modeling from Noisy Preference Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.263732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:71d6ce087df19eb3d390512fb062fdd3f79a590b2ad8435b873eaac6b4c3ee6c

Observation f3fc44e2-fba2-4481-ae1e-4045ff79a575 · outbound

This paper cites Confident learning: Estimating uncertainty in dataset labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference Confident learning: Estimating uncertainty in dataset labels

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.508583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:1f42409748cf610ffd08fa16ac1dde41be8c691e972e836a2072fd5bc052e42f

Observation 5bb3f878-5a10-4039-92e8-749572f89f0a · outbound

This paper cites Training language models to follow instructions with human feedback.

Optimal Transport for LLM Reward Modeling from Noisy Preference Training language models to follow instructions with human feedback

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.323448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:bffa3c3558ec9d266d888ab171e6a0a8a4c796cf0520b7a2e520bda841e690cd

Observation 27480ec9-280c-4101-b943-684f319a9ba6 · outbound

This paper cites Making deep neural networks robust to label noise: A loss correction approach.

Optimal Transport for LLM Reward Modeling from Noisy Preference Making deep neural networks robust to label noise: A loss correction approach

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.512604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:fc7200f1ce2731e975d75a96855b0fc1eef4fea16824a3459855c9313ade5629

Observation 82495d38-5e57-4ed7-9e81-8bf9fc70735b · outbound

This paper cites Qwen2.5 Technical Report.

Optimal Transport for LLM Reward Modeling from Noisy Preference Qwen2.5 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:46:07.391607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:fedc4b9d862fc0a210d46378f806460e45f99935cadd8a337ee41d4e0ce0cc6d

Observation 90cbcea2-1e7a-42f7-b75a-5a895857fa30 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.566431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:f174c34a99c47415335febb49e294bc40f797dc2f1eea44456dfc07bcffb085b

Observation 6d8d33a4-81de-4ad1-9a89-ae6c4222addb · outbound

This paper cites do anything now.

Optimal Transport for LLM Reward Modeling from Noisy Preference do anything now

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.420826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:e30a894ca0ec76cb99a023bd632e5876514e30eec62812160c52aa0aaa9c6c16

Observation 49b62e6f-fb1c-46fa-8eba-7434a6eb8367 · outbound

This paper cites Defining and characterizing reward gaming.

Optimal Transport for LLM Reward Modeling from Noisy Preference Defining and characterizing reward gaming

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.479808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:824817da2502b146f009a59ec162fb8933c907fd8920ef81be335856a658d655

Observation ef53027c-9cae-418d-a0b4-faa0060be3ad · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey.

Optimal Transport for LLM Reward Modeling from Noisy Preference Learning from noisy labels with deep neural networks: A survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.457995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:ed278865c4d82cf433ab8b67bf70b6ea22a204040eed9a7e90d4fe34206fa414

Observation 690463d6-b488-4dcb-9c89-913e6026e1eb · outbound

This paper cites Optimal transport for treatment effect estimation.

Optimal Transport for LLM Reward Modeling from Noisy Preference Optimal transport for treatment effect estimation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.380462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:18b7bb66e304e3a45ccac68f9bc70453cc5deb9fcd41cd1cc1a5082a5fb01e65

Observation bf09b776-103a-48e0-a424-24f1a5736e8d · outbound

This paper cites Unbiased recommender learning from implicit feedback via weakly supervised learning.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unbiased recommender learning from implicit feedback via weakly supervised learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.551737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:9dd521078d3cb5ef3822397a7fce4bc9d8359f376e0eb43e0cd787b47317a5da

Observation fa9b7335-b909-4624-a50e-1e7858217daf · outbound

This paper cites Optimal transport for time series imputation.

Optimal Transport for LLM Reward Modeling from Noisy Preference Optimal transport for time series imputation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.547992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4a40ce7c5e4e3949e78f475e34ee1f76eeac9dea6c6532e2922e63a738e91575

Observation 1cde68ed-2483-4110-a0c7-045269745308 · outbound

This paper cites \ epsilon\ -softmax: Approximating one-hot vectors for mitigating label noise.

Optimal Transport for LLM Reward Modeling from Noisy Preference \ epsilon\ -softmax: Approximating one-hot vectors for mitigating label noise

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.427914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:b370583078dbcef3ea80fd87d56e821555297ed49cd6df8ecbd8dd6a3f35200a

Observation 39429f42-4a47-4f38-bcef-e245d0a08f5d · outbound

This paper cites N., Egert, D., Delalleau, O., Scowcroft, J., Kant, N., Swope, A., et al.

Optimal Transport for LLM Reward Modeling from Noisy Preference N., Egert, D., Delalleau, O., Scowcroft, J., Kant, N., Swope, A., et al

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.504367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:222d4d34160a5130bc0a034fb2c07ae5b55416840fdc2baceb796e9345aaaa1a

Observation a8c1ba81-177c-49cb-bdd6-607068a1a0f9 · outbound

This paper cites To smooth or not? when label smoothing meets noisy labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference To smooth or not? when label smoothing meets noisy labels

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.835954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:ac9c09c3dcf7db69b963b44432a3d6c63e41ab0fcf6aa71f63dc0c9538f6ad2b

Observation 69762179-59c2-44c4-aa1c-274b792453c2 · outbound

This paper cites Revisiting consistency regularization for deep partial label learning.

Optimal Transport for LLM Reward Modeling from Noisy Preference Revisiting consistency regularization for deep partial label learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.820569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:8fc2babd524c0bc891a4f4edda995a6214e0c26dbc18f9f608d3cd12445e6d3c

Observation 5cc417ea-5eab-4237-8d06-971eb3b09cbc · outbound

This paper cites Robust early-learning: Hindering the memorization of noisy labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference Robust early-learning: Hindering the memorization of noisy labels

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.995766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:e6a9898026ff3f3271d001a307b90b800e98134d5daa02f64cceec088cb8ddac

Observation 12a8bc5e-ccbc-4779-a551-60bf578c2a5e · outbound

This paper cites Sample selection with uncertainty of losses for learning with noisy labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference Sample selection with uncertainty of losses for learning with noisy labels

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.079037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:cfcb36d6c8776e623a9067e12ab421e94cac6d308e5010c3f36b93a64e73aa99

Observation c3983a88-8982-4078-8051-652613b7ad61 · outbound

This paper cites A holistic view of label noise transition matrix in deep learning and beyond.

Optimal Transport for LLM Reward Modeling from Noisy Preference A holistic view of label noise transition matrix in deep learning and beyond

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.037225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:3d950b7a3b0f721dd1ab44dfac2f0f6b1f9467b0a593ecdf8d5e587c8cf0d033

Observation 783f012d-6b18-4a92-8536-3b8929c8f932 · outbound

This paper cites Early stopping against label noise without validation data.

Optimal Transport for LLM Reward Modeling from Noisy Preference Early stopping against label noise without validation data

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.044968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:1744adbf83dfc9649c08f1e1857ceb60df40256864e3671005d1c46e9356417d

Observation 16433f45-961f-42f4-b125-9730340138a5 · outbound

This paper cites Badlabel: A robust perspective on evaluating and enhancing label-noise learning.

Optimal Transport for LLM Reward Modeling from Noisy Preference Badlabel: A robust perspective on evaluating and enhancing label-noise learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.032882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:c94ea3f3009c967e2bdc378d1f8de206023db499175f485acea57e19f7529c1e

Observation de4d29f7-75a9-468d-a2ab-ff6d8f22b04b · outbound

This paper cites Clusterability as an alternative to anchor points when learning with noisy labels.

Optimal Transport for LLM Reward Modeling from Noisy Preference Clusterability as an alternative to anchor points when learning with noisy labels

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.041392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:8416ff6e921efb376b96460060dda498b343022be78d831068ead6938c16caed

Observation 53f47d05-e2f2-4e73-966f-85055a989313 · outbound

This paper cites 2025 , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2025 , volume=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.020582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:196dfa5ee9c0dfe1eba05f063afc9eb8997dd69a49128d10fceff77bde7b815d

Observation f28e22e7-ff85-4ad2-a729-de1ef19df9f3 · outbound

This paper cites 2025 , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2025 , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.934697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:cb491d582e5b592ed0fcd95c10dc93b5723117ce1808e5e137fdedce106c4e70

Observation 876ad7d0-f689-4f6b-815d-73c5704924b7 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.068338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:3e53b02a99f0427b5da0e34515af2dc63e7d2446ddfb5d9b9b1e9d5a3cfd9cbb

Observation 59e0a848-ebff-4328-855d-83f98cbf8f54 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.950527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:fe7a14e21ed9c3d0795f9cb7fe22170871883c27c6370f91cdd952d5ae168ece

Observation 79abf030-071d-4de9-978e-8cd025634896 · outbound

This paper cites 2024 , volume =.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2024 , volume =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.972309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:459b385c1aa12cd47063e284fc4921ecca0cbfbb22d579cbee648cf7b868fd46

Observation b6c70328-35b6-43a5-aea3-cc3c3e7a4195 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.976070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:27d131b1fb5a1deb661fa1e45310821d323f0e7b543e6d086c8e18db0451cdef

Observation e46c5483-f351-4285-9392-70fef11326d1 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.103969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:394b233cffe5f38cc4c2ae09990c8c4d88e4160b25e6e9bb015bdc5edbf64283

Observation 32aedf39-c492-4b91-b197-4587d01e4309 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.912462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:163ba85ea8d737344fda23dec654517ca36a6d0db301cc7259186d05a0aa9493

Observation 981581d7-5542-4c51-a231-b57e494e19ae · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.919744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:8726e7c25f65a8e09422a738e5c2e2f6acc60e071ae7ce95c4b3ec0b99499c1d

Observation c4ff76c9-e5d6-4f05-a3cb-095620561e63 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.905874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:947544111250e4e3c9ea52be79497e1bc22c2a11f38d0c549fd5c7499f45d5f4

Observation d103fea3-f0ec-4d0e-8d0f-ba2cce922ea6 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.024823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:e93fffbcf207c2cb763d87f3462e83080df67d321b4a08c5dab78210ef7ca283

Observation ec5677df-382f-4025-af09-07db5eb60e86 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.075650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:70c4ee8360be23a803c5fad4e1c6dfef16a31bbe5cce64e18229006721369e81

Observation de23f289-1db9-4e8f-9d24-cd61459859c0 · outbound

This paper cites 2025 , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2025 , volume=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.846998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:b3744ef309d82df96a4287f3f6457c5191188d4c853ac47835c3469c309a616e

Observation 8b480ca6-1e0b-4786-821c-2a73c0d39e50 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.999153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:8e06709df27e355f0f24c82089909a488ce3514168661461a4f7648b6c4f1cde

Observation 21bda3ce-05d4-4c99-bdd8-e161e2010b4e · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.402972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:fd40dffed7e15098f8a0b132efa7721b84ac947ce8d7e6c1e2afb97dbe0a1dff

Observation 1ea72eb2-198b-4579-9d35-82d122b9a4e7 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.417024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:e96fa73c86f369e3a342082394be16948f49e00ba51e0bdc26e7a2711de54e69

Observation 4979835e-57db-4a3a-8ef5-54182139fb42 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.337443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:7aed14c22992de033eab7ff16e91ab4067524f086031782bd75c012334e34fc5

Observation 1d84dee6-411c-4169-a503-c4bd1d1c856d · outbound

This paper cites 2024 , volume =.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2024 , volume =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.253024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:9330b74ece9174d4aede323b00167e729ab74a93843b66ace174d8b2a035e7e7

Observation a984db0a-ca60-4e7b-83f7-c55e87ffe28d · outbound

This paper cites 2024 , volume =.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2024 , volume =

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.028450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:5e19a133ff58603187a6240bd5d6f866b959a7817c8786f2bc49a5393a9d900c

Observation c7c2f511-4439-475d-b261-a51e536cf7a7 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.002258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:9280e8a1e404317c7e4942db7f8489ac2ff19fd7ac61bf8ccacd12a97e8c3138

Observation de5cd398-977a-4010-a6ee-137bb4bb6002 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.916183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:8a25a9dbec19ee70b34b34495482a30d0f331086419019073a667a26f70e16d5

Observation df470820-0467-42c0-97d2-d6763ce943e3 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.843374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:6468c5da2891c810c15b11237c6cd99dd5a95aafd7e73b04a9309be1ba147ea3

Observation ced60ec1-238b-4339-adfa-26a29ed93847 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.789951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:6d19ccba43dbc3f7c8895edf3ae7d10da0a017b3fd5aacee2caa2c29233577d3

Observation 885f2bf2-4fa3-4c96-a345-b5c29f4e65d8 · outbound

This paper cites 2024 , volume =.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2024 , volume =

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.793362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:f7169a877ab654d3be27b75efdc3afd95c86ad5c304a41097458b8cab844549c

Observation ca53b5cf-d86f-4616-af00-01da32764afc · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.801333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:836777ea1859c1e052bb97fed7fc43e65b1513393757eecc14120398b3004cec

Observation d5c42960-7e74-441d-9c8b-be4bb808475b · outbound

This paper cites 2024 , pages =.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2024 , pages =

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.284773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:373239bc9213d97e1a736e42f903419ca7f693642e66e3a3b77f6fe5e01ad16d

Observation c41fb3ac-6c2a-4a85-bf64-4592d62dccbb · outbound

This paper cites Kingma and Jimmy Ba , title =.

Optimal Transport for LLM Reward Modeling from Noisy Preference Kingma and Jimmy Ba , title =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.778525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:638250c98127cb3e27b4d476592091f93f4076ed40d8503b535c9603b7e43a05

Observation 530ee2fa-0f4a-49a9-b39b-870cde83f9be · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.782508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:aa64ec40e5c2c52c8f85f26905e859574d1ba43c3d64c60fd38b814a5a4866c9

Observation 5d80be32-47ee-40ed-bae9-2b705de689ed · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.773768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:aa9eaa3b413d8f44e1dc1965cf8f790118b91e9287e1911e866a98c5fad6bd8e

Observation 997b70e7-9f1c-4fdb-b3d5-23cd2570227c · outbound

This paper cites Nature , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Nature , volume=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.786286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:1563857e20385e1452cb1a250ad2e859340e3839118fc41d8dfd0b4d0dc13619

Observation 93fcd148-4025-4e2a-83b1-17e84581e232 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Optimal Transport for LLM Reward Modeling from Noisy Preference Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:46:07.329885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:58c781be4271cfeb1cebcf754a2cc5c1dfc3f5599a0388177b3be97fa6dbe70f

Observation 10086121-81b0-47ef-8ffa-25a83d342615 · outbound

This paper cites GPT-4 Technical Report.

Optimal Transport for LLM Reward Modeling from Noisy Preference GPT-4 Technical Report

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:46:07.424698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:92675429ace426564e2556746cb1ad7643ff6dae9f2a63fe4ce889eac844f657

Observation 265ee2aa-1c2e-42f2-9cf6-894082911c35 · outbound

This paper cites Advances in neural information processing systems , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Advances in neural information processing systems , volume=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.797496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:33689a00fb15c0fa32132d3bd456ca0ea11f2f3fef6f05a3f5d8b985e117791f

Observation 58f48342-85da-445e-b09a-cfcf4d9189fc · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.245235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:fd07d05b98921dd0eb533ac54e960fb7b7232bbde1cd21fcd8659aa1692eba24

Observation cf5240ce-a3fa-4e6d-a578-5f239f4acd22 · outbound

This paper cites the method of paired comparisons , author=.

Optimal Transport for LLM Reward Modeling from Noisy Preference the method of paired comparisons , author=

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.770042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4d852afb8fb3067ae768909d28634be318d4887209eeec45a52a01dc1fe0d017

Observation b236a05c-aaf0-44c6-a886-9d3f4a75ab98 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.241753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:576e717d34154cbcc9a9221ff3d5302879235c2603563df351d5d32fbfb1272a

Observation b8e3034a-293c-491e-9efc-181407af9ed7 · outbound

This paper cites Computational Linguistics , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Computational Linguistics , volume=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.115256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:d8d96419d3ab069a8e2ece41d3df26f5d1e17c84e3d04a13046425769db20ab1

Observation 875546cd-120b-48d0-aa20-7094924affae · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.009004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:94540c9519be4590a4d9a5bb01ce6ce021e07872b727b4106a62a6101142f0b2

Observation 01e7331f-40a8-43a3-a27d-8fcf94a6b084 · outbound

This paper cites On Symmetric Losses for Robust Policy Optimization with Noisy Preferences.

Optimal Transport for LLM Reward Modeling from Noisy Preference On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.420959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:cfa2d9c52633d12e60de3ca7655e964a24839cd60c22b884cee3df486f8b4980

Observation b33c64fe-ac10-45f0-abb1-ea8d20dfa0e4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Optimal Transport for LLM Reward Modeling from Noisy Preference Proximal Policy Optimization Algorithms

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:46:07.360705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:e7c02d688576f63d04a950f512383844cc1c27ce811b9083237b4656427a8913

Observation 850af893-dc20-46c2-bfee-0c26745db923 · outbound

This paper cites Group Sequence Policy Optimization.

Optimal Transport for LLM Reward Modeling from Noisy Preference Group Sequence Policy Optimization

Reference 89

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:46:07.445597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:c1de72696a57735f3b01f4e8f736652493c6f173a59555a854009b0c0bc27fe3

Observation b4b307cf-904f-412e-a100-7bf343200fc5 · outbound

This paper cites 2022 , booktitle = P_SIGKDD, pages =.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2022 , booktitle = P_SIGKDD, pages =

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.492434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:1c52314223d27cb486979da160608cbac37616b0d53a0ea5ddf1cd5948fc2472

Observation c682db53-2fb0-4aa5-bf5c-067eaef1a12d · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.436242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:16b2021f9e51f4aa1aad21e67b1628fed7aee6961ae3b64b710abf827d143a30

Observation 1593ee18-bf81-477e-ba92-8118dffdd86e · outbound

This paper cites Dual Unbiased Recommender Learning for Implicit Feedback , booktitle = P_SIGIR, pages =.

Optimal Transport for LLM Reward Modeling from Noisy Preference Dual Unbiased Recommender Learning for Implicit Feedback , booktitle = P_SIGIR, pages =

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.587194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4caa9515fc486bf7b7bea2d0e2030906ff52822a8e33a3bd5186c168db4d09d7

Observation 68a97bcd-0bea-4214-aaa0-64865a7bf2ff · outbound

This paper cites Biometrics , volume=.

Optimal Transport for LLM Reward Modeling from Noisy Preference Biometrics , volume=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.376395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:5a4b717bf048dcbe891a2969744015a04ee83aa6bb9b9117d55499d00b1a4f4a

Observation 4afdef64-8907-4922-8dc8-15781f082b87 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.282754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:3df5fddc3d385f162bcec01019c3ea5c771c63c1e12737a1f77bf5cb25d8a5cd

Observation 14c56590-710b-4ad2-ae6a-2b01cd5b2673 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.082483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:cf892c1dad8cdb1c8378e2ba67d323620e8d6eea4bb14c19aac046bc4f6a491b

Observation 2cbfb064-bedc-46c7-a750-c93b800b8d6f · outbound

This paper cites Large-scale Causal Approaches to Debiasing Post-click Conversion Rate Estimation with Multi-task Learning , booktitle = P_WWW, pages =.

Optimal Transport for LLM Reward Modeling from Noisy Preference Large-scale Causal Approaches to Debiasing Post-click Conversion Rate Estimation with Multi-task Learning , booktitle = P_WWW, pages =

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:16.968950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:d6e755455aa3c1333986ad9aee04bd0d145c561cbcac458c63841fb360a0a4e6

Observation 22cfa253-3592-42dc-ae5a-140b58eeb0b0 · outbound

This paper cites RecSys , pages=.

Optimal Transport for LLM Reward Modeling from Noisy Preference RecSys , pages=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.431850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:e910288edfa2f501d7ff9264353daf72835d3515d6bbbb964041b295573c9fb3

Observation d495223b-6a5d-4fbf-b844-86cc87ce830a · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.367637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:6b80dbe64487e3e19df14a57bac59f91165f1743114355ed743b4e6692675c75

Observation 4c92cdf5-f1d8-4bfd-a115-4c749e7b8d18 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.248102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:4530601b0db9d4f6941865cfc840a7b660ef7f952681f4b38e0c33bc38bdb647

Observation 6b602341-38e8-4cbd-a027-c2cabad07e39 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:40.277712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:c50d8a7bb9a52d8c01491a188246bf1e8f969358163b286b7982fa6cde9ec3b2

Observation ab3d96d1-338a-4c7b-aa64-8047afd8ba4e · outbound

This paper cites 2023 , publisher=.

Optimal Transport for LLM Reward Modeling from Noisy Preference 2023 , publisher=

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:40.296616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:42f34abe9b6cbd2a164c1c671cc9367508b51e03c7c0af5b6d0f05c324c72bcd

Observation e820ee61-d843-4990-bc28-6cd3a957430f · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 102

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.938622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:01768d7aeffd1b0337766d5f4d15b2382f2576438e22fb37858ee829a0e7fca6

Observation 1f6ad4a0-e004-4427-9019-92461382c008 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 103

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.061445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:9d393b7230ccb331bf5f1a4999056edc75f970cc6d674498909c69b627fb09bd

Observation bf91c744-9ee1-4645-9f1b-f7eebb817706 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 104

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.048166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:58a03488aaba959a33d6037cbd4870922d80de55f90ab0c45a69c30c7f37de60

Observation f31651d2-8230-4b9a-8c8b-11df6a179d4a · outbound

This paper cites The Concentration of Fractional Distances , journal = IEEE_J_KDE, volume =.

Optimal Transport for LLM Reward Modeling from Noisy Preference The Concentration of Fractional Distances , journal = IEEE_J_KDE, volume =

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:47:17.064922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:c6cb3159ef03e677ad8d90860b37114773b39ada385bc70b087eb29b6248dc12

Observation 0cd2fef9-8e84-44bd-a839-7e0515aa5d42 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 106

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.016520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:51d2921ea7bbcfb6135cd72b3feda4a427deda422d237d5bcbe2fbcf1ae91a10

Observation bdc6b8ff-f544-4ea5-b92a-a4b44641636c · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 107

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.058187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:a634e107f5b64dc32fea9c7e11a2badc4b257c79b0f498b2baae525df04363ee

Observation fe671794-9456-4e09-91ab-c41efaa69850 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 108

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.809079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:1fdb0525a094a04ed6ea04e55b1748e6f2d126899f53f2fb4a96e471704fe785

Observation b671b698-dfd6-4bf0-bfeb-3b0a2b687e9e · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 109

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.051561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:0e36bf939a529bea52749d4d0795dc44564efe05b24bc061a5fc5728680ca032

Observation 72d1cd9c-e24f-43dc-882a-31827d55030c · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 110

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.805567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:24a1e6acd14f47bf708328776d26481ad31bbf730f873eeb44f424337dc2ae6d

Observation c2e12d26-f13d-455b-96b9-b49484f24b87 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 111

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:16.812489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:30fd80028de6118c0f4c38f7cfc0d37b55ee8ff432efbda01c776a6a205049ed

Observation 3cc521be-a8ee-49d3-8ccf-8c301ea96dd8 · outbound

This paper cites an unresolved cited work.

Optimal Transport for LLM Reward Modeling from Noisy Preference Unresolved cited work

Reference 112

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:47:17.217756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:a994c573caf0b6e38593f3e0010249888e32e3023c0b188ac8d00e3da6aa861b

Pith citing papers

No inbound Pith citation observations are available.