Pith. sign in

Paper Citation Record · LEDGER

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2505.12843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12843 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:01.415106Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:28:18.050097Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T12:23:24.212503Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 687059e7-1487-40a1-974e-91027e8af4b3 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF A General Language Assistant as a Laboratory for Alignment

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.050427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.050427Z digest=sha256:1d11b43df2c548b9f56ff3b58ccff62d84dda6d31589219650d1d450d23cc538

Observation 1ca65930-4c38-49aa-9e68-b3412f374ba1 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.269453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.074161Z digest=sha256:a264b7bbd844c107bb5921addbca000580c8468032d3e5b3e38aa6d0849fd879

Observation f12f60ed-7165-4c6e-a46e-98f9aba21d7e · outbound

This paper cites Rank analysis of incomplete block designs: I.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Rank analysis of incomplete block designs: I

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.100010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.100010Z digest=sha256:dd2b8a0bd392e35034b7bddee1474500ae81ae351dc2b0fd28ae5c4287b1b0b0

Observation 8868765d-b6c9-4a2f-b86d-41ada71ae417 · outbound

This paper cites Noise contrastive alignment of language models with explicit rewards, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Noise contrastive alignment of language models with explicit rewards, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.242923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.115523Z digest=sha256:c92836ff9c570577c5af491a471ea554ddbb40112eb354fd27ddec707a976b88

Observation 68bc6e29-fd45-4f2f-b856-df39cac60120 · outbound

This paper cites ODIN: Disentangled reward mitigates hacking in RLHF.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF ODIN: Disentangled reward mitigates hacking in RLHF

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.227345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.120530Z digest=sha256:a6685e418f0eda2352f2d1324530e3533c3dfb71e5712bcb352d12122a47d636

Observation e1febd69-ddd5-41a2-9081-693e79f9f182 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.126627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.126627Z digest=sha256:bf2dc9c651b6673c058c5888b251e2a003ffcdda0d04fe435723f5e700844e9d

Observation cbd7e727-5072-40ec-88c6-d92c1175db1b · outbound

This paper cites Deepseek-v3 technical report, 2025.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deepseek-v3 technical report, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.197954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.133388Z digest=sha256:f226ec8d6c78a06c0acb2832a1c732480062eb042b4325e4bc17261ed85c64b8

Observation 1d399348-7ffb-40c2-bc44-ac4d26134666 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.138963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.138963Z digest=sha256:45b89a73c0cc28f436e2b080cc0eeebf49818d2eeec6e7076079ded938c5d5b8

Observation d96ae6cc-6e97-4c26-b659-3484f7806aa1 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.144235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.144235Z digest=sha256:a889271f4326c865ff7c591b7aeac166eafae677c96f04a5a48d43ac30b01b73

Observation 92b36bd0-66db-4942-a8bb-206cd3b71f28 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.149460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.149460Z digest=sha256:9ccc037d3aae16ba515fbb37e0c951fccfaa7034b480cf924901bb79a6ab546a

Observation 895ee927-9f17-4fa9-96fb-e185860aa0ce · outbound

This paper cites Helping or herding? re- ward model ensembles mitigate but do not eliminate reward hacking.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Helping or herding? re- ward model ensembles mitigate but do not eliminate reward hacking

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.170133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.156657Z digest=sha256:fa4b3d422e50594d1687549a67e7851bf2fd3712c64a44dcdb3f18d4076188a1

Observation 5175fce3-dbec-4130-828f-6b4db048fca5 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF KTO: Model Alignment as Prospect Theoretic Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.162942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.162942Z digest=sha256:8f57a9623bf8a680d4e420d098ca6bf6f5504bf311c288516dd73324f88772e7

Observation 00a8a2f2-ff85-4c42-92a9-3563c8715d8c · outbound

This paper cites Scaling laws for reward model overoptimization.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Scaling laws for reward model overoptimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.153446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.169483Z digest=sha256:fd684333a5600806657f8db484772cdffed6fefc3cc7d502e7df056a15929136

Observation 9f39062b-9580-4429-8b48-74aa5924b5a6 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF A general theoretical paradigm to understand learning from human preferences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.136000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.176821Z digest=sha256:b763eac86596c9706ee5e16275f6257249234036905f5f2eeee3205b1655bf0f

Observation 091fd5b7-8e92-4857-99b4-3fe00829aab4 · outbound

This paper cites The llama 3 herd of models, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF The llama 3 herd of models, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.183228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.183228Z digest=sha256:a4240b245a0f7437dee01a472ae32dad164baca985778a964a33dad8cc549897

Observation fd0117c1-6cf9-4226-8cb2-79708f860e0e · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.188382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.188382Z digest=sha256:2c06b3a38cfc79e03e146baf833daa98177184756c342c9d377daf4d92f2546b

Observation fe66a8c7-218f-4903-9363-d582253f066f · outbound

This paper cites Deep residual learning for im- age recognition.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deep residual learning for im- age recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.109230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.198616Z digest=sha256:30ad2bab26ccd866f87697274f233ebd9888dd821e10fe364bde82c06bb0ca4a

Observation 74341c34-490e-4297-a989-a09a27556a28 · outbound

This paper cites ORPO: Monolithic preference optimization without reference model.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF ORPO: Monolithic preference optimization without reference model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.089711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.205375Z digest=sha256:7b3584868a4a0f88e8c1354cd1e3aeb516ca22b20612f4e658f3da492674023d

Observation 461bde1b-68cd-415b-9a4b-068b4cb0c487 · outbound

This paper cites Post-hoc reward calibra- tion: A case study on length bias.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Post-hoc reward calibra- tion: A case study on length bias

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.070083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.210800Z digest=sha256:7ca2293367a800232c5d5bdc6c0d53e158a8434356a8a65a0dcad00f321bc24f

Observation 78169754-008c-4a47-93ab-923e4ca0c041 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.217132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.217132Z digest=sha256:60335106087ec8be828d098d5eaa5eee50360fdd9215d6458f5d5ba8fd336a94

Observation 021ab3bf-76ee-4e76-89d0-184714636678 · outbound

This paper cites Openassistant conversations -democratizing large language model alignment.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Openassistant conversations -democratizing large language model alignment

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.037894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.223306Z digest=sha256:9f271870cdf5016fc4acf368338bb1a2e9dd64fd9350f0dec1879410e9abb383

Observation 2d3a95c6-e250-4a0b-8ff2-59dfa125a0ef · outbound

This paper cites The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.230482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.230482Z digest=sha256:bb20c1515130aadd0f85a5f3f0fae6df16227d4a59f450bebd7104df08100750

Observation 73322ea2-4061-4ac9-aab0-9f0e9fb8f00b · outbound

This paper cites RRM: Robust reward model training mitigates reward hacking.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF RRM: Robust reward model training mitigates reward hacking

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:02.019581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.236592Z digest=sha256:af04b342400c243a0de10bfcfc4cb1cd9160186b94c92f7c3d34abbebe74d5e1

Observation ab8961fa-4a15-4bd6-b065-bbda41a69a15 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Simpo: Simple preference optimization with a reference-free reward

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.999843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.242997Z digest=sha256:8a44eeeb3fd4b0b0dbc15976f93b410d2c8a6a016cf66b0e28c543836243408a

Observation a758781c-ef87-4653-ab11-23cc8fff87c8 · outbound

This paper cites Expanding on what we missed with sycophancy, 2025.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Expanding on what we missed with sycophancy, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.981268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.248564Z digest=sha256:d8057a06ab6fbc89128ebb3f4c084e8be112ee4a82b39747f6243cf44a644bdf

Observation 77de0340-a9e3-4520-94d6-15af306c4b2d · outbound

This paper cites Gpt-4 technical report, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Gpt-4 technical report, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.964195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.253305Z digest=sha256:e38ce2633cf4013d8923bb597f9ec9f3df91d9d9e49cf26274c37ba94e7c7e4d

Observation 5cd313b5-52c4-4f27-a8c4-4934d7f0dab3 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.271309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.271309Z digest=sha256:ad354f95a656493ba7d553968fe7553b1e647cfaa9f5ec89a403af2c6964f486

Observation 0a0d94a4-fffc-452d-9a16-c2a02c146196 · outbound

This paper cites Reward gaming in conditional text generation.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Reward gaming in conditional text generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.937446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.276692Z digest=sha256:3ff7334f2d6907fb5d91a341f939e1475eac356e6052c0db71dffc27c7d8cb4c

Observation 2627ea87-679b-48fb-893d-a7fc06ed9222 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.283312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.283312Z digest=sha256:b5387cb729c5bfccffcfb1d526fbe3831b976e52e4efb7cb2e30b15be990b56f

Observation d92e127b-ad5e-4f4c-a8c4-19e72c3bb138 · outbound

This paper cites Notes on regression and inheritance in the case of two parents.Proceedings of the Royal Society of London, 58:240–242, 1895.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Notes on regression and inheritance in the case of two parents.Proceedings of the Royal Society of London, 58:240–242, 1895

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.922834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.290347Z digest=sha256:d6b915be5803a3f0c096dd17cd27f543c26a53823de1b30d778a5476335efbfa

Observation 37e87cc3-4853-45bc-90ff-18db27b3ab96 · outbound

This paper cites Qwen2.5 technical report, 2025.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Qwen2.5 technical report, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.906666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.296424Z digest=sha256:2a32cbc0dd9a57af962615d2ab459c0ddc4557c93b4abdf32356ae2e90425517

Observation 9a55eee8-8a33-4a5b-aa8e-d25da946b6d6 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Direct preference optimization: Your language model is secretly a reward model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.300971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.300971Z digest=sha256:d69f0a40a440ea3c5b3846c9e14fa034b90895ec5c65a8ba1a5ee05b6bfe3275

Observation 7b0b67e2-78f0-45f2-9c4e-f6777a4f5f11 · outbound

This paper cites Warp: On the benefits of weight averaged rewarded policies, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Warp: On the benefits of weight averaged rewarded policies, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.307466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.307466Z digest=sha256:70a7d0103d48cc72cbc775092fb0efef4c754127b1f27e3d33c6486ff92fd56e

Observation b5c25c3d-c1a6-4e36-a4ea-024a25c267d7 · outbound

This paper cites Warm: On the benefits of weight averaged reward models, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Warm: On the benefits of weight averaged reward models, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.867011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.313692Z digest=sha256:e4174ffe189021c842b992879bbb6d2929ff7dca9534c329830bf24152e1004a

Observation d6da3a41-b5bf-438e-a7b8-50564ac0918e · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.850062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.319597Z digest=sha256:5acd6cba25e3dec3128606753a644ddb63dc359878928bd2b7efd79fadcb0f1a

Observation 0b795c10-9d72-4afb-9121-51894a263dde · outbound

This paper cites Offline regularised reinforcement learning for large language models alignment, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Offline regularised reinforcement learning for large language models alignment, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.834087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.324999Z digest=sha256:0acec90944f6cd13c4363b0771bcdbbfaff8568f7ae64ed2dcff6b5c62f4d3bb

Observation 18fb0222-e90a-43ce-9ef6-c3ceb3a0cfba · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.329871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.329871Z digest=sha256:835b456fe50fecd6fb053f8333e7cb7db7dfd57617500c0da44a5744a5d4749c

Observation 94b50415-f300-4540-aae8-db9f96f76aee · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF BOND: Aligning LLMs with Best-of-N Distillation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.335192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.335192Z digest=sha256:b6e502660d2afe0dd64d34395623e3bff7c2f41d7e029e0cf60a3fec7ad76f3c

Observation 1a72128c-6975-459e-8d5c-fba75b26620a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.340582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.340582Z digest=sha256:0078775a27202820fa6febf907f40a278933aec732716fe7d695590e8d85c2ad

Observation bbd4fdcf-3d6d-4e65-96e9-9abeffdcf289 · outbound

This paper cites Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.817025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.346657Z digest=sha256:2333ea1d176ac9262df2543479dd85fef6ff522ebdb1efb07950971e1d8e21b2

Observation f0ccf053-ec92-44ca-8847-175f6aa45da3 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF A Long Way to Go: Investigating Length Correlations in RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.360484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.360484Z digest=sha256:6e6d26be6b5876a4914a5ed09b3b58469be531cb729d37e44459580517993944

Observation 6d44421d-76c1-4dc6-94a5-d28e21e8aac4 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.365962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.365962Z digest=sha256:783de9f1d040a097a2f69dd544b876f6618ce2f8d0e01926400086963f2f1cd7

Observation c6973f78-33e7-4b84-83db-5d1a703a0c78 · outbound

This paper cites Learning to summarize with human feedback.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Learning to summarize with human feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.370763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.370763Z digest=sha256:714dbb9022cb5a13524b733574835c1df157234d3c0902353e5b56a66fb9dcd4

Observation d4a46a08-4386-4d1e-86f5-7e78004053c8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.778601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.376369Z digest=sha256:0c3fd0cb59bd38522e6c249e19c5fc7e97245f752797d51a28037ad54eeae540

Observation 75dc0f5f-8422-4368-853e-d3a0f1184e20 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.381925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.381925Z digest=sha256:97e2b1d3a1b6a9a0fce68582a148e411ccff42129324e33469873957300f9740

Observation ff50780c-1bc8-4d80-824d-ddc9add43235 · outbound

This paper cites Attention is all you need.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Attention is all you need

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.386796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.386796Z digest=sha256:3335ddd8418ab6c248012deee9426c8d03c8c2c2faae4dd2cd631c2e236cba2b

Observation 7a4b8528-1e2e-42d9-8532-c0aef22b84d4 · outbound

This paper cites Reward hacking in reinforcement learning, Nov 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Reward hacking in reinforcement learning, Nov 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.738195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.391674Z digest=sha256:5c2b8303f9e7cd4326ed7ef3a22c6938cb0de1a99e7092d7a984f0e41ecd19a1

Observation 08881065-5386-42bb-9196-9ca184a1626c · outbound

This paper cites Qwen2 technical report, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Qwen2 technical report, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.397242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.397242Z digest=sha256:85f87b0f82b67461ff846d8078c1bbcf60d56528d52ef466edee511c87492756

Observation fe08b45c-ef18-4e96-8f0e-097c11063e79 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.707100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.402730Z digest=sha256:cb9255cb9c91e833ed1be9bb695852a6400560a8800908cfb691c0d9da5b6c33

Observation 947656c9-9a9e-41d4-9a11-c51b5d8c3a15 · outbound

This paper cites From lists to emojis: How format bias affects model alignment, 2024.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF From lists to emojis: How format bias affects model alignment, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.686042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.408429Z digest=sha256:ca0219a7f5490264fcfdc3c7c849d3318176814390bf79f378ae4246bf74df57

Observation b184816d-d7cd-4090-838b-07fa64a1387e · outbound

This paper cites Fine-tuning language models from human preferences, 2020.URL https://arxiv.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF Fine-tuning language models from human preferences, 2020.URL https://arxiv

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:01.668211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:31:01.415106Z digest=sha256:65836fb38f29320cbfc6daf1a075656b29d4c5ff1a5f441a0be4cac0a7b4ca45

Pith citing papers

Observation 83119876-bf91-4161-9c30-bee51499f1ec · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-06-25T02:17:29.855075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:6249c3acecc52445febe1262ba2c1729d6b6d9328d3bbbeadc35cc6944e9c8f7

Observation 01dc5012-628e-4f96-8630-5ab77989de2c · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-25T02:17:29.855075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:eaca9e68b4a2e0e15e4d2f3459e97a5f1a9eae77a17f11c42f548e74639b038f

Observation a1c70b58-51db-4535-98ad-f9d2cd75863d · inbound

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure cites this paper.

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:23:24.213633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T12:22:06.635622Z digest=sha256:e3fcc59a1e621db58160a84a8733acf9d0335ea72c39934ba6b08e389871964c

Observation 12f1842d-a9bb-4ac9-a585-9b1f71f32c38 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Bias Fitting to Mitigate Length Bias of Reward Model in RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:18.050097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:18.050097Z digest=sha256:53e457a4620c874d2714be2e086d077b10fee2e3719ebbc8325c6c3a6075d537