Pith. sign in

Paper Citation Record · LEDGER

Debiasing Online Preference Learning via Preference Feature Preservation

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.11098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11098 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:32.728004Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97cd3f7b-566b-455b-9cdd-a3f0b49084f3 · outbound

This paper cites online" 'onlinestring :=.

Debiasing Online Preference Learning via Preference Feature Preservation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.520994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.520994Z digest=sha256:c52688be66152e8fe31b7c3996746c911f6c3a0cb60bb1d5e8461e60c96be577

Observation c393ac49-6827-468f-9471-6b4e1b397011 · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.526128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.526128Z digest=sha256:95006bfb493fc125dbbedd6a1e9e5fff97a51d4d31dc72a793925c9b5248892d

Observation b14c64b4-b5fe-4978-a3d5-703f650aecdc · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.531806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.531806Z digest=sha256:61277a0775c393f250a5e167794ae231515806ef039f0f70f354c53b0fc61f95

Observation a67aad3a-e58a-4e61-a615-7a741689fdc0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.532490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.536525Z digest=sha256:a5a3fcf8b9ed9d98342725b2d5b3ca13c6b7faa4031b6a837a689000a4286aa8

Observation 2acc3515-d81f-4cc2-8512-0b46352bdb5c · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.518687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.540968Z digest=sha256:da5d82acda26cec845264982aa94574cf63ccf6a386841766d28cb6d2a9b7e8a

Observation 44a430db-54fa-440b-b31f-c752a6d5620a · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation A General Language Assistant as a Laboratory for Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.545630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.545630Z digest=sha256:5551dc3a2a9a7977a81c03f39aff19edcfa187fd4341dfa13bc0be2260b2d9dd

Observation 54f27b40-29ad-4e25-8375-2ab3cc997f5a · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.504015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.550254Z digest=sha256:d03a40486b9315005d00f5404807d7aa55b398e225b85a8163f6e98257bb8457

Observation 4638ed7e-9013-4d13-a010-f0d005282a97 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.554671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.554671Z digest=sha256:845e38a00fecfd95f316940e182fc4311222d12ba87ca5e1eea9689b9ecb9fe8

Observation 1fa1e298-330f-4b76-91a0-94a7d1af9242 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.479058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.558585Z digest=sha256:0b0ada8d858792368b521e23f28ce5eb2e47fa5533594ed85fdfe65fe449bd47

Observation 39cd5008-b9f1-45cb-b3e0-b8c2368a4106 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.466212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.563093Z digest=sha256:9bdc879df8b1b6517935b946cb44b9483398a7c2428b159f8265bbf12de76244

Observation 15bfb751-8677-430a-ba9e-c97906086dd8 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.567554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.567554Z digest=sha256:576fa39927644dda5aa1f64993da72a4ee33ab23cd72bfdef8c6ca7ed987c1ab

Observation 37b32f76-af8f-48b6-844a-3ffdeb628b8b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.452443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.572066Z digest=sha256:7d36981ad71534a5f4b2a489a9950689fb02745bed96ab609f35f35aeebca7a3

Observation 14915f7f-eb95-403e-91cc-6af505725a08 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Debiasing Online Preference Learning via Preference Feature Preservation Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.576314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.576314Z digest=sha256:f0f63b5f60b7f53894130a017d45e29baf2c65e617b2d25c9459330ff5f8165d

Observation bbe05557-40df-4a4b-a544-a963a7750947 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation RLHF Workflow: From Reward Modeling to Online RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.580792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.580792Z digest=sha256:8221849b99293e4b2c08dbfa6c6fa19d897d44a9ca0b978ad72b569fd59622a3

Observation 67bf7084-0174-4fcb-99f5-a281574c56aa · outbound

This paper cites The Llama 3 Herd of Models.

Debiasing Online Preference Learning via Preference Feature Preservation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.585102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.585102Z digest=sha256:e82b751db0938765993c11be15f21af3e9e6dade5e70e054e330c8934d265afd

Observation 0209bffc-0395-4431-9978-bde3421de69e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Debiasing Online Preference Learning via Preference Feature Preservation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.589245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.589245Z digest=sha256:b74a1cb0f857de17fa828b4045e691bf43a8e718f0bfbf628470d3e0354d85d2

Observation ec8188f9-2433-4f22-a2bf-cdf944df81a2 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.438711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.593625Z digest=sha256:7d89bf7f4ed3796356b9a4c1b413b0949f25680c3e7251a8b38b9505cefd2bb4

Observation 4a59ccb9-e276-465a-8636-d63bbbbe37c0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.425259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.598590Z digest=sha256:71115330e8e20f5ef1149e5f8a4089e2427d10ead6c992910f4c75a206e34721

Observation af5dfd71-9369-4733-ac9b-992386de1f21 · outbound

This paper cites Mistral 7B.

Debiasing Online Preference Learning via Preference Feature Preservation Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.603244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.603244Z digest=sha256:6a057a321c95364b06fd9fe6395f9f4bf35c5ac3e90d1b4927e5c815959faea5

Observation df7810e6-6911-4b53-9744-df9bc024dc31 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.412157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.607485Z digest=sha256:20969529da8b7e7edb1ca1b4ed1665c5c2c42fb05870ca46b623688c67d9817e

Observation 231ed846-a0a5-4267-8112-b3975ad68ce0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.399238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.611469Z digest=sha256:30248ec953253db770f2d045c0813bf25d5774f8440ecedf56eee166fcf05320

Observation b44231d1-45fd-4c50-8cbb-88ba9410d503 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.385631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.615612Z digest=sha256:38be0687aba85ab8ccafb359dc18bf826d71bed16de9d7c0a7c6b14574607e17

Observation 2c603c55-2ae5-4eba-ab22-10ea4b4bced8 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.372321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.619403Z digest=sha256:f29411fffa9d1713ef01b1a11659ef3c6d8b8c5be8af605aad50f4c00d08c396

Observation 3955dbab-d44d-4b18-82f6-e7a3f4bea961 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.358683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.623578Z digest=sha256:3c9a5f9eaee366c7b0d31eba63236b79dd67dba00de2b41785ff722d0bd3d865

Observation 94169d60-021c-4f72-8000-9e67f334bd5d · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.343913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.627509Z digest=sha256:76a4e3003ee849cd4f98921e592b0d91847c5fc8e93f50ed8e52d75214020f21

Observation 7c27b1ca-6f27-474e-8f27-021ee75597fa · outbound

This paper cites Dissecting Human and LLM Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Dissecting Human and LLM Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.631391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.631391Z digest=sha256:18d1ceccd440213f15cc2126806b05df0c17e4b3829ade3fcdf35ee6ff30a2e0

Observation 4cd640aa-6fef-45cd-ab5a-98e22d37fa91 · outbound

This paper cites Decoupled Weight Decay Regularization.

Debiasing Online Preference Learning via Preference Feature Preservation Decoupled Weight Decay Regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.635828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.635828Z digest=sha256:012bd1d2a96aa3449bf05aa0451019835d7c6b7810690a089cfaa10739277380

Observation 6c8ed455-6fc1-4c28-9a79-47d7443ba863 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.330752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.639794Z digest=sha256:7277d14363a517c749b1e8d9c1dbf16355fbb6653f44b4da18aab2077a4ddb19

Observation 0050e4e5-2c19-42df-b80a-5aefc823c192 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 29

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:05:33.018292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.643899Z digest=sha256:c5d7c28f970de5968ead16fcf012f64a9db2dd61691b8f8650b4692af740f78f

Observation 134dfb4e-b7b2-4554-a56a-0ec9ccb4fb01 · outbound

This paper cites GPT-4 Technical Report.

Debiasing Online Preference Learning via Preference Feature Preservation GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.647804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.647804Z digest=sha256:12d9f3fb8673731de35265796baff751cffd898c606ce3308c099b75d8758fe3

Observation 0d800e66-12e2-4516-a0b7-70f26436693b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.316887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.651800Z digest=sha256:b638fc6450687f3013c8a02305d13e63fad2a25ef71808bdc822c755e7db5133

Observation c16bedea-7899-4537-b161-7e33e1cf8054 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.656381Z digest=sha256:814d7fe5ba25221f8106e2832c34d3777760acdb7728cc00014099b22df4a787

Observation e3c64ba8-a585-434f-8d09-81744b4e0b12 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.289403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.660320Z digest=sha256:54d5f33135d50f60a123cdbfda63f0bb0ae37b5c4719de7b3ba27a17861c3d5b

Observation 7ccfae8f-eb6d-4547-99b4-4b9439e73479 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.275333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.664289Z digest=sha256:337db0d06ccf39c2c44675460c0442807d39fd211af0c22e2bb95782dbb6562f

Observation 0449bcc5-cd22-43c8-8a98-5a9be9064e05 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.262059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.668431Z digest=sha256:c876a6885bc28b27f9e6faa59a1f7995c80a51682136647454a0435a713f7734

Observation 0fc2bec0-5495-40c1-a5e9-18cfb280fb3c · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.672178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.672178Z digest=sha256:2ad4caac93dc0f0f5a40bcb2b90306f02a18d2ef3f1d5b4541af482f97f3f2cf

Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.676825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.676825Z digest=sha256:b1f0da4f0dd2af7ecf26e858fdd6e0cc02943de99810d035c5ff83ca1c84277b

Observation c1ca1cb8-b1c0-4130-afcc-3099e44ece02 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.248054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.680958Z digest=sha256:210874feaf588537c674af9bcc5874aff7b7d7bc819b9f6df705602ca4cf031a

Observation 0b4f36e2-831a-4a94-b610-5e9d3d917d0f · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.232005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.685213Z digest=sha256:8cbe70b35c268edf8bd2974815a5e4b8a6e9e2f87ef4281bb2d3bac6df31130a

Observation 362a54b7-755f-4ab6-b9d8-4a615b1153db · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Debiasing Online Preference Learning via Preference Feature Preservation Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.689590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.689590Z digest=sha256:89d6dd54be02b3f4f750ddef1958d2ef24379f872c73d9029e938109957d22d1

Observation 242ee8b0-5754-4c86-9e0e-6a9315e892ae · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Zephyr: Direct Distillation of LM Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.693883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.693883Z digest=sha256:0fa0ab0a77be973d0fa99f2a4f0af4fa445fa6b8122589eff82a9ecfe2c61097

Observation fdd9d55c-10bc-41fe-9be2-aeac5ce84b3b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.218139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.698505Z digest=sha256:6bba75817418a12556d66ec4cc126aa9dd853fdfbbda45ec4a486f00aa113e1a

Observation 9192bb21-5105-4608-9b76-1f00beb25dc1 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.202246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.703255Z digest=sha256:9055f9325bfdec0b424c6a1a0407964c28fc85bd1cbc0be91e435e675af10914

Observation 0d0b9669-c8b2-4b83-8629-aa35d4e4397f · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Self-Play Preference Optimization for Language Model Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.707537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.707537Z digest=sha256:c8f1f2c869680d7925e4c2ce92691bffd218c5847ec5f69dd5709a00596304f5

Observation 944ea925-9df0-422d-bede-ecfb6cef2c23 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.188133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.711756Z digest=sha256:6d751971f8fb823c2cc3b7aa5825bdd8e96f849c41755568ccc3aebb899956e3

Observation 8a00e02c-ef7c-4cd6-b98f-5b30ad389630 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Debiasing Online Preference Learning via Preference Feature Preservation Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.715738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.715738Z digest=sha256:d2f2bdb86208f51b4075fe782e59b477493ded27b87afba2130ee2b105dc8ae7

Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.720004Z digest=sha256:6ae6599b95203d2d11384f92a545cd0b9652902985e8879ab6c68d410484bbbf

Observation 0bea02bf-8970-4335-b3c0-29668a3e6023 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.174460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.724090Z digest=sha256:f3a990713f24d6f35d226052f579ee4b5d1e5a09a717f534fd286a7b2fd9ade7

Observation 385a7655-4429-4b52-86ac-4b787c317a94 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Fine-Tuning Language Models from Human Preferences

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.728004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.728004Z digest=sha256:904861ef93020f2a8ba9d8781176e6fc6d9b52822d2c7d84155ded5936ac9d19

Pith citing papers

No inbound Pith citation observations are available.