Pith. sign in

Paper Citation Record · LEDGER

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

As of 13 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2411.12843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12843 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:15:47.969050Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e161b2c-34e3-4133-a110-a5e585e1f317 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.761360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.761360Z digest=sha256:f403e0dd22803cd9cf7e00f4f4ea56e4dea65d2b37071fa831b8e0d845e74272

Observation a25c7f74-9667-4f33-b6fb-4c84c54ce535 · outbound

This paper cites write newline.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.767020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.767020Z digest=sha256:5d474d3def30722258c73a967953545c0bf4485809608f86a5676c0cd6bf20ed

Observation 081e54fb-3acf-418c-be61-60cf8a58c732 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.570039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.771511Z digest=sha256:021b7dcdd764000707d398075922fbc512ef713820e94b1b39e828b2b3d8aac2

Observation 3691b473-c2fe-40e0-9867-c3f02393bc96 · outbound

This paper cites Direct Preference Optimization with an Offset.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Direct Preference Optimization with an Offset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.775855Z digest=sha256:68b696cf183f3030ecb9406bf015b2cad29e74d1936a5b43febb70fbc6695b3f

Observation ec021049-0eab-49c7-b308-445dbea10f05 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.780377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.780377Z digest=sha256:e9ff5ffe6239e311d7c9e2ba827319af328731c775714b5991de23fb46fe6219

Observation bf53b892-14f2-425c-a9a3-ab687e1ea5ec · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.560838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.784982Z digest=sha256:ddf42b23e5f6b8960a6f4daee1972dce77f6c88cfd623a6260f8f22dcb75ed48

Observation b6c3f1fc-d1d5-437e-89cb-e94123dd7724 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.551619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.788433Z digest=sha256:cb67d88f3075d873c79dcbce6d3553ed51f08208fc4252a9731d2d92e566d067

Observation 40253885-eda4-42bb-897d-6536a1eb4854 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.793228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.793228Z digest=sha256:87a9b365355819d17ea3227c7e31639a4c3bbb0a71412746ec192e752c60e645

Observation 6d254da8-a3c7-41f1-97e3-76f4752b190d · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.797841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.797841Z digest=sha256:4592342b722c9d6b82e8fa3fed6703664457365ba6118c00b05e93305443d791

Observation 2ef17b6b-f8b7-4189-ab8b-30825ccca466 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.535395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.801861Z digest=sha256:1d1782ecf125f57f18b8b506c2a34b5875b36a1dae0698cc3d234d26f2a17661

Observation 3d06af9e-fbb3-4650-ac11-200bed482f14 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.805929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.805929Z digest=sha256:cbc1ca73c16ccf56e67be472c45214939f44f298c5b0bf1295d504e9b4f55c9e

Observation 70a8b507-2a26-45ac-9373-33f3a7df35a0 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.810658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.810658Z digest=sha256:d613a73f015e638cbd57e33fbee6370ce54a0e0c01c41a2288ec94ee2506968c

Observation d935d9db-872b-4b7f-96cc-af7637e00f58 · outbound

This paper cites u rnkranz, Eyke H \.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd u rnkranz, Eyke H \

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:15:48.523675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.814466Z digest=sha256:ca3c913576b9d97774547636a1a3206b2917aa4354c4840da60466240dcbe7f1

Observation 1a394c47-48e2-4352-b8af-e43e3388513e · outbound

This paper cites On the Weaknesses of Reinforcement Learning for Neural Machine Translation.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd On the Weaknesses of Reinforcement Learning for Neural Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.818566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.818566Z digest=sha256:2b4bde0740a543ca547576e6f8f9c74da7fb896fbaa9daa948021c78247bdf97

Observation 152af994-d48c-4e32-975f-2259477455ae · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.512901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.823139Z digest=sha256:2e246d0c17da5162e64994755550f77450633b2c731cb2475274e51cb79a77c1

Observation b4e678f4-cd4b-47b1-be74-b7f539c50a42 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.827057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.827057Z digest=sha256:74282278062d1f3e781f8f57ebea96c4a47ad5882163812dacdecd103f3b6f42

Observation c3dc8528-77f9-4e5b-b3b4-fd185d323166 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.502602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.831232Z digest=sha256:aa6e12959ac9166f07ea4122d5af7787a9daa5c3e5dffb19c1f47fc0b4d1f4df

Observation b83db4eb-5179-401b-a927-dd8a7d208015 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.835813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.835813Z digest=sha256:2bece1e1bfd9948a95efbe8c2198ef6bd266728e6573c8a6f6c3d82bf11a7c5f

Observation d77b52b3-ffcf-4168-ada9-63c5d20a9e18 · outbound

This paper cites AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.840044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.840044Z digest=sha256:9c275adda427b1a68f5d4a0372b72c44dd050b1f1375d31173141c7fd2ca5f01

Observation 3af6c8b8-c15d-43ce-8191-56798173b0c4 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd KTO: Model Alignment as Prospect Theoretic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.848222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.848222Z digest=sha256:7d21a8223b46fbc56864ff621287aa23705b9bb6dcfbf6a1b60af6ad628a9e0c

Observation b902db48-cb28-4ae0-84b9-70ae807cc43f · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.852311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.852311Z digest=sha256:03c7d57a473ea69ab5521cb6505420924d1e9b4711c0e8bab8ef59a19283551c

Observation 940f65f2-140c-471d-b43c-2e6c68a55b92 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Distilling the Knowledge in a Neural Network

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.856138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.856138Z digest=sha256:038cf008c309b029e44bbb41d78cfecdb1707679b6e797fd9853f80439ac7a9a

Observation 9a96e8ad-c492-4abe-bda4-17336cf459b9 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd AI Alignment: A Comprehensive Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.859637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.859637Z digest=sha256:dd3ccb5241709055b86243dbc55b703e0e821b6da0c769f5b2237d09d63771ac

Observation 2561aac9-bb22-4ecd-b954-2056a00fcb8d · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.863766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.863766Z digest=sha256:edfa1bf27557e49143e28dc31e575a25462330ee092378cc2e2f826050399da7

Observation 940af647-dbf0-44da-84ef-cbae7ac48c00 · outbound

This paper cites Smith, Hannaneh Hajishirzi.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Smith, Hannaneh Hajishirzi

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:15:48.492158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.867933Z digest=sha256:1d873e554031d90992d9dc947a88607ebb163606b645f3cacd90a5279d7853d2

Observation 5ff87c1e-2a71-461b-9b7c-9b5836edc896 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.872874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.872874Z digest=sha256:4c9a84b3a10b5a73025b3469bdff775a568e9095ffb861af47b6fbd1ab4ed7ec

Observation 92c73e22-97ef-45f7-855b-f8d2ae6cf105 · outbound

This paper cites Reward Learning From Preference With Ties.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Reward Learning From Preference With Ties

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.876946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.876946Z digest=sha256:afa6925b62f69371a5cc18b0128223c29281f063f28c205775701ab172fbf8a4

Observation 0bb642c8-3734-47cb-8b39-7bd02e952bdb · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.880969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.880969Z digest=sha256:dd5bfbf50c732a86e3b5cc85158da3b33864f3f51eec89a6f5c3fce011faf85d

Observation c5b7fa9b-491e-42b9-b450-131bf43bd4af · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Statistical Rejection Sampling Improves Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.885158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.885158Z digest=sha256:03058807c43dcbad95fe56f99cb87bfd33de65cdc0f05c0b49fc62ce511f80ff

Observation 87abdf8f-3f4b-4d97-bd42-e8bc9a804c22 · outbound

This paper cites The Llama 3 Herd of Models.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.889376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.889376Z digest=sha256:5e155c8e44718efcc9a3db857f176494ab1f0dcbf56409324ebf5094c3d10258

Observation 93e065f2-3805-4137-a6b4-2d950aae6d9a · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.481699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.893254Z digest=sha256:4e1233322372f3ddc24e662aaadb25e3a7fae4629d0f428b1bdedae6897a6392

Observation d7f4663a-1ec0-45f7-b058-b44373011881 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.898069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.898069Z digest=sha256:585bc096fd0b072214f2d41303897ba245aa3db8ea4863deab9aed088af08ba2

Observation 6c9964f0-f602-4602-9447-f84738a543bf · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.464869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.902599Z digest=sha256:c5994d58a76d7772c199d1ca4cad8a5c24ca35c03da35f392f0414e15a456e6f

Observation 68a67b04-bd87-4310-8196-92aa9276446d · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.905917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.905917Z digest=sha256:67d2f92874cdd541a809233a91b5d3a6a0dae978443d003765360627dad7194c

Observation a5988650-d0e3-4e86-b80f-e694e68a69ba · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.448474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.909873Z digest=sha256:780c9582fbcffa06ae2eb432fea2e8f7c5fe40909574ae0579572a05f0739195

Observation d3072f0c-8627-4e49-a563-cc20e1536aff · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.438091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.913664Z digest=sha256:58160dd9e32d76a86b1c7a04bc4f3babf150faf72e5e0425a448a210b3931614

Observation b0922306-e35a-4b56-a400-cf09f76f5d87 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.428510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.917263Z digest=sha256:16e76644dd887536a5966f2582aee1940d5b817270d87c197803628e484d3df8

Observation 3f8bd9b5-5eb3-496b-a319-81820c7e1d00 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Gemma 2: Improving Open Language Models at a Practical Size

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.921096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.921096Z digest=sha256:39fd32a7602e67b2454c3e8f4fbc3d689b50d3e1af3aceea575159cc50157c1d

Observation 250473cb-99de-4577-8adf-89c68256c17d · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.925208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.925208Z digest=sha256:97879bdc60aba2e11195a6d5bb59b84ba0131fa0412418977024474aea97b87f

Observation 0c1c3830-a08e-4dee-aea8-068e398d2bfe · outbound

This paper cites Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.929415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.929415Z digest=sha256:0f272bdfc4bf270aee2ba7343bfae17f0519a59c972d3f7734eb9e006464b554

Observation 6586546e-2a46-48dc-8144-e1e528e25074 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.933909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.933909Z digest=sha256:d38ded98731d4d1d7461b9c3b113b561839e5efc896e9875dcf112bf7cac6447

Observation 5e5792f0-9ebf-48ac-bfaf-dcd252318112 · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd HelpSteer2: Open-source dataset for training top-performing reward models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.938224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.938224Z digest=sha256:0d3434b0b707610bc5f8703c57882545344b890b054a03cb3d19eabc0f8da024

Observation 2cb5f4f7-c7cd-4489-a6f2-2162d6cc3936 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.942379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.942379Z digest=sha256:2038d477512a6f486558830350fdcd626f49d22ae224f7af19fd552c0fc72812

Observation 74005b54-252c-498a-b0b3-0efe579b79e0 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.417739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.946180Z digest=sha256:315eae12f050bef5c1cfbfe3c8782b2fcb8ff7d3fca65f185c5199dba706bcad

Observation eec568d8-294d-4e6f-ab4a-d7e74b41319c · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.406353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.950140Z digest=sha256:8ba1b986d2d5549c91c81516c8c8b63acc8cb610f8975d21512f6c7034f47520

Observation 14559811-d16a-449b-a098-a619feddcf6f · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.394128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.954036Z digest=sha256:b4554d20fa9f1e912bd48dac7ae1be49ee4c1167ef2078b5b5ad33842dba62b0

Observation ba846ef6-2e07-4e29-a80a-bb825c1e5a78 · outbound

This paper cites Token-level Direct Preference Optimization.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Token-level Direct Preference Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.958038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.958038Z digest=sha256:068fc4e313a2758159c79a290e859cce0a7d21c6d70177402e95e720545686d4

Observation 9a74b634-c129-40ca-b1e4-d533e99c7df5 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.961445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.961445Z digest=sha256:ec040a41de6ce24ec1c826b3300b64f76358bd0d9126a8e3ae6f7e27a6b52bf5

Observation b4cbc0b9-fd3d-459d-8dee-dc44595f53ad · outbound

This paper cites Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.965644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.965644Z digest=sha256:4a362f51db895461e95ed80500a3bffe232db5f3cb79dc101a43cdad4b5e643e

Observation 9af073f4-65d2-49b9-b1f3-9963811425cb · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Fine-Tuning Language Models from Human Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.969050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.969050Z digest=sha256:f2da427a32f52fb884815a2f3f31d07fcf08ea1138ac986f58ca739cb01f61e7

Pith citing papers

No inbound Pith citation observations are available.