Pith. sign in

Paper Citation Record · LEDGER

Debiasing Online Preference Learning via Preference Feature Preservation

As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.11098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11098 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:05:32.728004Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 97cd3f7b-566b-455b-9cdd-a3f0b49084f3 · outbound

This paper cites online" 'onlinestring :=.

Debiasing Online Preference Learning via Preference Feature Preservation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.520994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.520994Z digest=sha256:5190c57ac7c6b3c3fd772698bf8a90c971628574b2f039112023e86fa9a83df8

Observation c393ac49-6827-468f-9471-6b4e1b397011 · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.526128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.526128Z digest=sha256:c61f2c6780676419cbb514cc18d7a4abff172560b52acab5220aa5e355a01aeb

Observation b14c64b4-b5fe-4978-a3d5-703f650aecdc · outbound

This paper cites write newline.

Debiasing Online Preference Learning via Preference Feature Preservation write newline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.531806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.531806Z digest=sha256:de51cd2f469db36b2ddec98f19bdf8d2f08b075f69f3a3ab5e6866f5cc32701a

Observation a67aad3a-e58a-4e61-a615-7a741689fdc0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.532490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.536525Z digest=sha256:419524d213fb25dbb1fb9e7eda03dd4d369e4e019ce149e2a6de366323bc4c16

Observation 2acc3515-d81f-4cc2-8512-0b46352bdb5c · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.518687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.540968Z digest=sha256:d234d313348600a27a276337dcda6139a4c2933474b7671bb7684e91926913dd

Observation 44a430db-54fa-440b-b31f-c752a6d5620a · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation A General Language Assistant as a Laboratory for Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.545630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.545630Z digest=sha256:d8fa4e424bed8f6af752283ec5b582e3770e8c040d41a7d308197dc17e97a157

Observation 54f27b40-29ad-4e25-8375-2ab3cc997f5a · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.504015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.550254Z digest=sha256:9efaa1b98f65f2ec0176f33fdfe86760c0c957174ff1b46ea9abde05cb70aecf

Observation 4638ed7e-9013-4d13-a010-f0d005282a97 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.554671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.554671Z digest=sha256:9047333b64ca68f92ecf8f26e9e2c816d6dd2fe8c5598e58ed34175a595e41fe

Observation 1fa1e298-330f-4b76-91a0-94a7d1af9242 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.479058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.558585Z digest=sha256:e22bf95190af4bb30f245df2afbc6401b10c56fb48f2e5d25616f4e4c8132940

Observation 39cd5008-b9f1-45cb-b3e0-b8c2368a4106 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.466212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.563093Z digest=sha256:66acc105c430c8cf73eaa406ce47ada7830370df8bf5fdfe396cd67dd20262cd

Observation 15bfb751-8677-430a-ba9e-c97906086dd8 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.567554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.567554Z digest=sha256:0b0fa8eb63f82a48fa6133434808820b9f5698ee6827fca97adccb2e14b0c99f

Observation 37b32f76-af8f-48b6-844a-3ffdeb628b8b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.452443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.572066Z digest=sha256:a8317c8148b9eb21e7bcf9ccb1f75145ba786bbcb4f25725379c582f75694bfe

Observation 14915f7f-eb95-403e-91cc-6af505725a08 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Debiasing Online Preference Learning via Preference Feature Preservation Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.576314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.576314Z digest=sha256:dc4da98d602b90a05d93ff9bf45a6fde5be847415bb3d554b775681fc508ff4e

Observation bbe05557-40df-4a4b-a544-a963a7750947 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation RLHF Workflow: From Reward Modeling to Online RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.580792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.580792Z digest=sha256:707dd41413b399e541a23198f73ebbc2be698a5b35c7c6c160e9540734f162ce

Observation 67bf7084-0174-4fcb-99f5-a281574c56aa · outbound

This paper cites The Llama 3 Herd of Models.

Debiasing Online Preference Learning via Preference Feature Preservation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.585102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.585102Z digest=sha256:9e6dca26f159262af7a764e070d2959162ea1d96e4296438d93faabff8afa857

Observation 0209bffc-0395-4431-9978-bde3421de69e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Debiasing Online Preference Learning via Preference Feature Preservation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.589245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.589245Z digest=sha256:9eacd7b567779496720576f4433dd1760041bc545d47dc72e5cbc4b4929e2021

Observation ec8188f9-2433-4f22-a2bf-cdf944df81a2 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.438711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.593625Z digest=sha256:49b06dfe39b7952ee9a6e254f566ae80ae19d18466aef2276b743bd38ea14216

Observation 4a59ccb9-e276-465a-8636-d63bbbbe37c0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.425259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.598590Z digest=sha256:7f5236226ebf13f9d84129d213d6c291fa979cfec108581c988e44551aad707d

Observation af5dfd71-9369-4733-ac9b-992386de1f21 · outbound

This paper cites Mistral 7B.

Debiasing Online Preference Learning via Preference Feature Preservation Mistral 7B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.603244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.603244Z digest=sha256:0ca2d1f4cb9e5ff6309cb8369004cf0879cc557aeba1651eef60ff5fd909bcb4

Observation df7810e6-6911-4b53-9744-df9bc024dc31 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.412157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.607485Z digest=sha256:229ef1bb0dc63d78dc241ba93b59ee1c20bcd0fe1464ee927c4e6329efdda2a2

Observation 231ed846-a0a5-4267-8112-b3975ad68ce0 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.399238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.611469Z digest=sha256:07a6f27386d431a04a0cf4bd242b8bf7ba63ea79b1142e9c0729d6e1756c5ed0

Observation b44231d1-45fd-4c50-8cbb-88ba9410d503 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.385631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.615612Z digest=sha256:25844f66c0b350a1581878b49a257b712c2f0d65b3a06f990f28117caf79f479

Observation 2c603c55-2ae5-4eba-ab22-10ea4b4bced8 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.372321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.619403Z digest=sha256:1f48acc823074f0e21fff38487a23139e1dc4e98f0a5b81b734adb0f41f5fcdd

Observation 3955dbab-d44d-4b18-82f6-e7a3f4bea961 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.358683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.623578Z digest=sha256:1e2ef5ae7592242ffd1ee3cb9c1d9cce59c1a4d23bd15cd318c7a98fe4fc7998

Observation 94169d60-021c-4f72-8000-9e67f334bd5d · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.343913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.627509Z digest=sha256:1b26b2b3d48875edf88eb4b997f1c647a5f8cd6c5505c1ee70e85f0d89a3fb72

Observation 7c27b1ca-6f27-474e-8f27-021ee75597fa · outbound

This paper cites Dissecting Human and LLM Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Dissecting Human and LLM Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.631391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.631391Z digest=sha256:490513749bad61fd38d742474479c9bf3a10b931168fab5139ef628d81638b88

Observation 4cd640aa-6fef-45cd-ab5a-98e22d37fa91 · outbound

This paper cites Decoupled Weight Decay Regularization.

Debiasing Online Preference Learning via Preference Feature Preservation Decoupled Weight Decay Regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.635828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.635828Z digest=sha256:0b701f21e70488d990c5680a2c6fb1be06af57f8708aff1a0266157d1f99bb7c

Observation 6c8ed455-6fc1-4c28-9a79-47d7443ba863 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.330752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.639794Z digest=sha256:5ca6d14cbb777e3a2c2877854ebd9269c2fa176a88845cb91dff24773e85ff46

Observation 0050e4e5-2c19-42df-b80a-5aefc823c192 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 29

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:05:33.018292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.643899Z digest=sha256:06f3f22b076e651053c6954bec4cc79d810c2b949a825bfb6e41a3ed702fcebf

Observation 134dfb4e-b7b2-4554-a56a-0ec9ccb4fb01 · outbound

This paper cites GPT-4 Technical Report.

Debiasing Online Preference Learning via Preference Feature Preservation GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.647804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.647804Z digest=sha256:3208e380dab13c86198cda2f3999f513a1889b160cecc1d067b963ac6905ab9e

Observation 0d800e66-12e2-4516-a0b7-70f26436693b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.316887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.651800Z digest=sha256:3a5aef28380d0a76f18f17a2b5d07427941a2e667894f133cc784e3e940f12e4

Observation c16bedea-7899-4537-b161-7e33e1cf8054 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.656381Z digest=sha256:5e2243fca2a31133dfb3a35d76027e04a4dd5cbd1c2cd6f7e260ca0263067c86

Observation e3c64ba8-a585-434f-8d09-81744b4e0b12 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.289403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.660320Z digest=sha256:44bb32ac3851fc444213d2749a0c8b96af452b4fe19b928f1876f763d48a7e27

Observation 7ccfae8f-eb6d-4547-99b4-4b9439e73479 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.275333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.664289Z digest=sha256:07f6e91dd4e815e1c92c52660f88c1afe276f4e81f7fe8c62a82dfac420dcebc

Observation 0449bcc5-cd22-43c8-8a98-5a9be9064e05 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.262059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.668431Z digest=sha256:7c866308e87aea4cd7e9f91ffc16be10adb9e45fbb6bee4cf76c877284b73dc5

Observation 0fc2bec0-5495-40c1-a5e9-18cfb280fb3c · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.672178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.672178Z digest=sha256:502789769870a667c367151c6bac678a659ec391ee655cc78e5cfbf228e10b95

Observation a3844b5c-fc83-4018-b6a9-5f94054c4f26 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Debiasing Online Preference Learning via Preference Feature Preservation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.676825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.676825Z digest=sha256:57d3b0629e3297d774247f902f16b0a7e85b88a005f50721510d0977c326cb58

Observation c1ca1cb8-b1c0-4130-afcc-3099e44ece02 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.248054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.680958Z digest=sha256:d4392eed1e04966548e2ece825236f272d42a1919391901dd43211e365ef5328

Observation 0b4f36e2-831a-4a94-b610-5e9d3d917d0f · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.232005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.685213Z digest=sha256:a45dfe24c9c70075e82b8263c22988c0de5b3b6b0b5031644f3d367756987737

Observation 362a54b7-755f-4ab6-b9d8-4a615b1153db · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Debiasing Online Preference Learning via Preference Feature Preservation Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.689590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.689590Z digest=sha256:8db841835572c17b6735dd84954a554fcd3dbf93b03357d1453779f1e176db71

Observation 242ee8b0-5754-4c86-9e0e-6a9315e892ae · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Zephyr: Direct Distillation of LM Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.693883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.693883Z digest=sha256:e7dcea0fe4fb0f99f667bd8fb95aa54eac69c9ce2e0491b52961b5e5bcf88ab2

Observation fdd9d55c-10bc-41fe-9be2-aeac5ce84b3b · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.218139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.698505Z digest=sha256:28c369d5a3dcb05b8802885dbc737680358c7f9003a4106629f84177952fbcc0

Observation 9192bb21-5105-4608-9b76-1f00beb25dc1 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.202246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.703255Z digest=sha256:30b247674641392218a03382bb94cf4eb53920836a1cbe67d13853fbc1fe1ebe

Observation 0d0b9669-c8b2-4b83-8629-aa35d4e4397f · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Debiasing Online Preference Learning via Preference Feature Preservation Self-Play Preference Optimization for Language Model Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.707537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.707537Z digest=sha256:ebbdb18546e03988d75a6cf55446b5fed88af48b4cec6ffc9ec9e6ead8e233ac

Observation 944ea925-9df0-422d-bede-ecfb6cef2c23 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.188133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.711756Z digest=sha256:721b595d766acccbcae7e7e6f3c1c02121a2b3a15592309ae387cbf8ea9cc9eb

Observation 8a00e02c-ef7c-4cd6-b98f-5b30ad389630 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Debiasing Online Preference Learning via Preference Feature Preservation Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.715738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.715738Z digest=sha256:4b83c442b1c9eea538a9b88830fb910627afa880b9b1e43fccdbbfab60580576

Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.720004Z digest=sha256:dcaf4ce801f2950da45b325c2360148d7123d3c1e6f9952c65734a6ebad46feb

Observation 0bea02bf-8970-4335-b3c0-29668a3e6023 · outbound

This paper cites an unresolved cited work.

Debiasing Online Preference Learning via Preference Feature Preservation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:05:33.174460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:05:32.724090Z digest=sha256:95c26bb8b8caa138fcb7e720476e58e8106922389c1766d676fc5b27b2409244

Observation 385a7655-4429-4b52-86ac-4b787c317a94 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Debiasing Online Preference Learning via Preference Feature Preservation Fine-Tuning Language Models from Human Preferences

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.728004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.728004Z digest=sha256:06f2fc5c1a600fb0c0ca6275b622e45bd5968dcee8e320a7f9c72f76949c7e5c

Pith citing papers

No inbound Pith citation observations are available.