Pith. sign in

Paper Citation Record · LEDGER

Debiasing Online Preference Learning via Preference Feature Preservation

As of 23 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.11098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11098 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:32.728004Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97cd3f7b-566b-455b-9cdd-a3f0b49084f3 · outbound

This paper cites online" 'onlinestring :=.

Debiasing Online Preference Learning via Preference Feature Preservation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.520994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.520994Z digest=sha256:6e8571bf61b99faefb6aa9c9dfa77158a72003e96e314ee04fd019b117c7b2ab

Observation c393ac49-6827-468f-9471-6b4e1b397011 · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.526128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.526128Z digest=sha256:dc66b8c3ef34eb941bf3c642afcbfe5eb53b12662c815d9801d904652d3bf4df

Observation b14c64b4-b5fe-4978-a3d5-703f650aecdc · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.531806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.531806Z digest=sha256:ab57493e7d8d09c9cc493de1dcdf919dc1aac1db27b8b8e1a0209bc90b489e66

Observation a67aad3a-e58a-4e61-a615-7a741689fdc0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.532490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.536525Z digest=sha256:245e0b463700ce939798af31caabc2366f00d9920b3db28451658b255823f8bb

Observation 2acc3515-d81f-4cc2-8512-0b46352bdb5c · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.518687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.540968Z digest=sha256:f5ddb69ce0a97dcc4e3f67ec64b2cab093179d654cc37c5dcc971b163fe23395

Observation 44a430db-54fa-440b-b31f-c752a6d5620a · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation A General Language Assistant as a Laboratory for Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.545630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.545630Z digest=sha256:5a6bbcd61766868014605e558018d2ff6fb53b0d59a5054d435ff7605ad7efcb

Observation 54f27b40-29ad-4e25-8375-2ab3cc997f5a · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.504015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.550254Z digest=sha256:b1d9f3fc820d0a18e0a5ea62a237e78e0d2bb66a35eb0097dcebaebac5a074cc

Observation 4638ed7e-9013-4d13-a010-f0d005282a97 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.554671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.554671Z digest=sha256:16a6aeb503bba83140c459a7229a0a1ba6e8cd98f23854cf681856f63ea28085

Observation 1fa1e298-330f-4b76-91a0-94a7d1af9242 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.479058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.558585Z digest=sha256:23be30588a02cf8829fa95816bc786b9309185ef5d3e9e2faba894403c4ec06f

Observation 39cd5008-b9f1-45cb-b3e0-b8c2368a4106 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.466212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.563093Z digest=sha256:99c4ad59c3e96825646ae0993c072a9ab6c1066d16ebbd44a217940161f601c3

Observation 15bfb751-8677-430a-ba9e-c97906086dd8 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.567554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.567554Z digest=sha256:2f133d0b0b619ba7f645924bc9b056bcc6ef8894b89c6817d79b74f78958fa39

Observation 37b32f76-af8f-48b6-844a-3ffdeb628b8b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.452443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.572066Z digest=sha256:1aea501e818479e630817fa36410b3deb50d29bdad67a63abeb1875fe76d3b6f

Observation 14915f7f-eb95-403e-91cc-6af505725a08 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Debiasing Online Preference Learning via Preference Feature Preservation Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.576314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.576314Z digest=sha256:673895512e4da87d0d168036d07bacc6f087d4006b63881175678c06ee208e76

Observation bbe05557-40df-4a4b-a544-a963a7750947 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation RLHF Workflow: From Reward Modeling to Online RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.580792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.580792Z digest=sha256:411319edd146ffe9b33951edc44d5f5645316a8e70238e39c7e10893d657d083

Observation 67bf7084-0174-4fcb-99f5-a281574c56aa · outbound

This paper cites The Llama 3 Herd of Models.

Debiasing Online Preference Learning via Preference Feature Preservation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.585102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.585102Z digest=sha256:ada181d970c0bed99b39c7d27820badc0b5e0273b6a27dd2131230f7ba342d70

Observation 0209bffc-0395-4431-9978-bde3421de69e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Debiasing Online Preference Learning via Preference Feature Preservation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.589245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.589245Z digest=sha256:db2cdc44243fa0861c65b6e24e74d54befdbf89d5052c3b579801136685b2362

Observation ec8188f9-2433-4f22-a2bf-cdf944df81a2 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.438711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.593625Z digest=sha256:43d5b7904768f7c51370f6477946d368ed10633b210771f7c0ad47b5981dc4f7

Observation 4a59ccb9-e276-465a-8636-d63bbbbe37c0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.425259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.598590Z digest=sha256:f2f0d30050cbe415b20b1bd0638fc1571949d93f8d4774e75c587e8d35cbf50a

Observation af5dfd71-9369-4733-ac9b-992386de1f21 · outbound

This paper cites Mistral 7B.

Debiasing Online Preference Learning via Preference Feature Preservation Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.603244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.603244Z digest=sha256:e0d0b99c673dcc6addcef571c9598053b5f1363d21e30818ac8840e999a62bd7

Observation df7810e6-6911-4b53-9744-df9bc024dc31 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.412157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.607485Z digest=sha256:8ff0252f24d833297b1ecd1acf6e70e65356d125fe07d404df52a194ed69c1c6

Observation 231ed846-a0a5-4267-8112-b3975ad68ce0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.399238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.611469Z digest=sha256:a73ac3a9d14b20542409253d222ae8ea849f2f4ce7259722e812b4e9c951cbf2

Observation b44231d1-45fd-4c50-8cbb-88ba9410d503 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.385631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.615612Z digest=sha256:6f520cb21058b1bbe8469d001715775d861cb8e86693792001b573fefaeecca6

Observation 2c603c55-2ae5-4eba-ab22-10ea4b4bced8 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.372321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.619403Z digest=sha256:11547c5e242beecd0f10b0cdab6c029cf2e6ff6506e0e11ccf6d6e81f44d447f

Observation 3955dbab-d44d-4b18-82f6-e7a3f4bea961 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.358683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.623578Z digest=sha256:37cd4a996f1c856750b6a9658a9cccb1d312293e43d28683301b1a4d94cc78e2

Observation 94169d60-021c-4f72-8000-9e67f334bd5d · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.343913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.627509Z digest=sha256:955337164b904578200f95d0c6dff8a1ab5b7623d458ddc007bc3e4531c4387a

Observation 7c27b1ca-6f27-474e-8f27-021ee75597fa · outbound

This paper cites Dissecting Human and LLM Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Dissecting Human and LLM Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.631391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.631391Z digest=sha256:b10158aa0321f2e1455a84729f2098f6bffa207304b797d08200fce5116bce65

Observation 4cd640aa-6fef-45cd-ab5a-98e22d37fa91 · outbound

This paper cites Decoupled Weight Decay Regularization.

Debiasing Online Preference Learning via Preference Feature Preservation Decoupled Weight Decay Regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.635828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.635828Z digest=sha256:a06f94adaf584df548473b90e09cee3244aa954582abbfbce56f705cb845f1e6

Observation 6c8ed455-6fc1-4c28-9a79-47d7443ba863 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.330752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.639794Z digest=sha256:6eef4724ae6b02a71cc4ecb6cff7496a839b5a674312938f9f46923a29f46f63

Observation 0050e4e5-2c19-42df-b80a-5aefc823c192 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 29

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:05:33.018292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.643899Z digest=sha256:d0eaaf1b8d34a793cf5f2b83e95ae0431beec15e7a98e3ed46c550d106dcc731

Observation 134dfb4e-b7b2-4554-a56a-0ec9ccb4fb01 · outbound

This paper cites GPT-4 Technical Report.

Debiasing Online Preference Learning via Preference Feature Preservation GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.647804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.647804Z digest=sha256:4af13f2fbb7efce3bba44cd87e2824f2d62693a5b3ad1d270b719fff4e13316e

Observation 0d800e66-12e2-4516-a0b7-70f26436693b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.316887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.651800Z digest=sha256:76de18b75fb3c663df218e0b294aaa3b678bf83ed4578be660561919bd2cb5e0

Observation c16bedea-7899-4537-b161-7e33e1cf8054 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.656381Z digest=sha256:3e19bc586bdd5eadbe6bf812e7304706206620b8cb68ad40d5228a6e620fee02

Observation e3c64ba8-a585-434f-8d09-81744b4e0b12 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.289403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.660320Z digest=sha256:37db253de06b595e62058b61622af76155f58a98bbf5d09dd78f4b3a6c057f06

Observation 7ccfae8f-eb6d-4547-99b4-4b9439e73479 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.275333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.664289Z digest=sha256:f1c334c05126d6706dfedb6372f2a9387a62d7734336f99508e4703fc1ed0002

Observation 0449bcc5-cd22-43c8-8a98-5a9be9064e05 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.262059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.668431Z digest=sha256:50023f4bdcdfdf7ad7d699f718822c9c1da656dca986e19c36e4c359b96f6d1f

Observation 0fc2bec0-5495-40c1-a5e9-18cfb280fb3c · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.672178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.672178Z digest=sha256:eae883e177de4b4a5703e5d14a72c6df98616175697cff03504fdf9f6ff97038

Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.676825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.676825Z digest=sha256:0feb4dd4ec8440790f4d6504567a54d8a4b4e69b10ea35fe6d6d60c467448ac5

Observation c1ca1cb8-b1c0-4130-afcc-3099e44ece02 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.248054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.680958Z digest=sha256:df09c924771abc8f7e7a3ace5b2efc43c8fa02e25bfc8c9cd723edfb7cbcb88e

Observation 0b4f36e2-831a-4a94-b610-5e9d3d917d0f · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.232005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.685213Z digest=sha256:fcfe919aac0503ca66852db44eeae09004f641c376e850d413d2ca5291058973

Observation 362a54b7-755f-4ab6-b9d8-4a615b1153db · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Debiasing Online Preference Learning via Preference Feature Preservation Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.689590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.689590Z digest=sha256:ab9be90d99b127cd98c684477d04060ed57f9ad36d00dc6dca616b3d246d0f7e

Observation 242ee8b0-5754-4c86-9e0e-6a9315e892ae · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Zephyr: Direct Distillation of LM Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.693883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.693883Z digest=sha256:34f27101b279dc03cdf07f84fbc98bb443a77b21e034705fae324bc301f9e6d8

Observation fdd9d55c-10bc-41fe-9be2-aeac5ce84b3b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.218139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.698505Z digest=sha256:0687fc613fec9d44ab7754a60deda3f2e8a008a6ba2b10695ac0636df250abd3

Observation 9192bb21-5105-4608-9b76-1f00beb25dc1 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.202246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.703255Z digest=sha256:b496d838abf078d278c7ed0095651e6a6e097c46496af9b635674764c9ac218d

Observation 0d0b9669-c8b2-4b83-8629-aa35d4e4397f · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Self-Play Preference Optimization for Language Model Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.707537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.707537Z digest=sha256:2fe1a0d8cb97ddf715a193036c0021dacb06a425b891edf21d1b0d6723289fb2

Observation 944ea925-9df0-422d-bede-ecfb6cef2c23 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.188133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.711756Z digest=sha256:0c682493041337d87d6760f50ab63b72d57e428feffdd5fe95387f3c7730833d

Observation 8a00e02c-ef7c-4cd6-b98f-5b30ad389630 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Debiasing Online Preference Learning via Preference Feature Preservation Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.715738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.715738Z digest=sha256:c06b34b06268b52edf9105da0601d07f772850f238cf2b8e339a403218ddf2ff

Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.720004Z digest=sha256:9959af126ac64e2dd2548722ba701f40a5ae2d4177c070b8f18658faf1021dc4

Observation 0bea02bf-8970-4335-b3c0-29668a3e6023 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.174460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.724090Z digest=sha256:767e133ba717beea67c545583c5e07d02dea9396292682c2680498c40bb5db5e

Observation 385a7655-4429-4b52-86ac-4b787c317a94 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Fine-Tuning Language Models from Human Preferences

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.728004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.728004Z digest=sha256:2c72f8f7e4c5c356b67b8a900a6534bc7dfb449d43130dce14b2ffe0b7d85ee2

Pith citing papers

No inbound Pith citation observations are available.