Pith. sign in

Paper Citation Record · LEDGER

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models

As of 13 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2501.06248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06248 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:31:45.734231Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63b0bb0d-2444-4081-ac61-41292b01466b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.586146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.586146Z digest=sha256:28743d40fca4a86ebd6d7f4a2d5b2aac5d00345e71f0a7ccba93a7e658996932

Observation 69553556-f3f6-4317-9aeb-0f9ea8d6161c · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models On the Opportunities and Risks of Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.591475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.591475Z digest=sha256:a57874ebef70561e8608aa1f185eb2bd4600d10aefb87410c6e26c91ed22a7ec

Observation 056526bc-d7af-4d73-a5af-30a5c2f08331 · outbound

This paper cites Language Models are Few-Shot Learners.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.596208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.596208Z digest=sha256:0ca3924a1094329959d87422c244963c9ceb4d1cfca7c7c9b175b2118b19b315

Observation e3b3b00e-561a-4824-ad5b-31f4003490f7 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.602059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.602059Z digest=sha256:7723dc00a549a0ed64a03e6cac312b0571965af0953b741ccb4f7eeae3363eb5

Observation 817eda52-d0f5-431c-bd61-25f2726e6f93 · outbound

This paper cites Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.607085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.607085Z digest=sha256:47dbd715f3b56ff5f897e445a4c31872d28901397f072bb89e25273bde2a7cb0

Observation 3abef2ae-4105-4fe2-bd5a-372bbf6d1b1c · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.612012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.612012Z digest=sha256:ce80ed05293fc4e438d48a2df941bcca42844241c114f909fe50ba421adb4fe7

Observation 0ef1db20-cc05-4cec-9123-ba7a076ec3f8 · outbound

This paper cites Attention Flows are Shapley Value Explanations.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Attention Flows are Shapley Value Explanations

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:46.036410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.617402Z digest=sha256:4234ceaeeed307d1a1aa9b9dea73c385a335c06c371cc4babbac697519b30a1a

Observation d43c383f-e8ab-48e3-9731-411802b1b3a6 · outbound

This paper cites Axioms for AI Alignment from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Axioms for AI Alignment from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.621875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.621875Z digest=sha256:3dd622f88931af0ac2570cdcbd050eb089b54d1220fb67a5d669493d5a97d754

Observation 1c32782a-3e91-42a1-9c00-e763f65a879d · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.626337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.626337Z digest=sha256:fc31b8327375acf4c8d4ba4359b0bdf7b1a6ffee76ba45a21d6b90f78b9c92b3

Observation 12272422-9d9d-4b21-a3ac-c2ff4ce161fb · outbound

This paper cites Steering Language Models with Game-Theoretic Solvers.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Steering Language Models with Game-Theoretic Solvers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.630739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.630739Z digest=sha256:804421163c642b2d727339fa605499abcb6aac410a517d115e75315a29411a21

Observation 24db5438-2529-4b7c-8f0c-ccc38c6f2f56 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Improving alignment of dialogue agents via targeted human judgements

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.635195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.635195Z digest=sha256:9875dad549316ef21ea1cf25e6182656606dcd73f54fd006b12e1082280ac61f

Observation 6080f8c3-ea81-4161-ad10-30f11ebf0a5e · outbound

This paper cites The Consensus Game: Language Model Generation via Equilibrium Search.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models The Consensus Game: Language Model Generation via Equilibrium Search

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.639950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.639950Z digest=sha256:91e209a13db316df8a1d99613ccde0c316b72ddd0967f27a1cfbbbb6fccef412

Observation c3d976a8-3a21-4061-9a45-7ad1402a7f05 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.242306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.644489Z digest=sha256:e8d72a9d546a4166e2d85fc5bd924f95cf02c8412f8c43df92424737a58f408c

Observation 23e0c1d8-1691-480e-baa4-8a0aa3aa6c1a · outbound

This paper cites Confronting Reward Model Overoptimization with Constrained RLHF.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.648887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.648887Z digest=sha256:b92029adb5bc7672b4d4139387201ad8b87d292ca34d93139caaab083ac899d2

Observation 4cfdc89b-b94b-4f18-8c90-3004031e06b8 · outbound

This paper cites Nash Learning from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Nash Learning from Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.654182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.654182Z digest=sha256:552e5ade2ea0b1cfc9c5ba58e683b7da7301ec3bbf08d3197a62a7d25745c222

Observation 1b4998d3-5525-4a81-abfd-67e7a4ba4e99 · outbound

This paper cites Learning Social Welfare Functions.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Learning Social Welfare Functions

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:31:45.906234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.658803Z digest=sha256:b6de3845afe765e256c323f563e21c5e461d36d6d39bbaa6bed7726cbb5718cc

Observation 847d409e-a360-4a39-83a3-f8fdf8999ddb · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.227446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.663300Z digest=sha256:853c2552ad9f7ed13b553b1cc27fa3f1a843cbbd5fa418b3e1dc23214763d40f

Observation 771fc87a-285d-4986-bfda-c994e177d059 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.212247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.667488Z digest=sha256:c61b6cb4dcd6b2ffac9080a919866d0d46769eb1db71584a1fd4ebeac369ea1b

Observation b1bd732e-37d6-44c9-aa74-bade1ffe93d1 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.197705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.672008Z digest=sha256:0d2d818b85e6d3e96a1be68578ece1d8af075fc666c96fc83457fe51b788b192

Observation ddb1ab5c-24e8-42e3-bc45-a5ee8f733e34 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.183274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.676572Z digest=sha256:db4dc3bed10e303b3d757b45010d102011b3baee06d63bbfcb100ebb9c51f86b

Observation 65e807e5-2b58-4273-b5a1-8043d6d25af1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Proximal Policy Optimization Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.681223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.681223Z digest=sha256:2811b6e9a8cc5bb41fe67b3b3e8a5ff9c8e360a7399b34b768c66ddc2520be36

Observation 571b9802-8a55-474e-81c0-c2e5a4ec0f20 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Long Way to Go: Investigating Length Correlations in RLHF

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.685491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.685491Z digest=sha256:2596bb314befc3b65c348986e403ce26575d10327bb15533417274514f520804

Observation b86165bb-7e5f-458a-a718-31a4d8330e34 · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.690199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.690199Z digest=sha256:371fd43959fec38a5c680186e8d5f311be660ead59da1c40e39a55b6a1e7f007

Observation e699ef17-cf16-4f60-a1e9-a2a258580c43 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.694507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.694507Z digest=sha256:04423712ae7ec8a74fa2b70af77d4bda7bc2357f92b28d71eea1c8dc10ffa5b0

Observation aab78833-519c-4606-8f87-c64c2f219b4b · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.698903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.698903Z digest=sha256:974f544c975113e2d7cd535de943bfd7976e8a7dd3c8839b4e07802a857ec3a9

Observation e46d02ab-af9a-4a41-91e5-acf244c16eeb · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:31:46.158280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T21:31:45.703389Z digest=sha256:361353139720a6c860548e4638386be69b25962f77e668ac9727853d8b4c494e

Observation 97accfb3-d797-48ae-9bd5-3cb04c4287d0 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.707647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.707647Z digest=sha256:d1f0a9ae747f1e5b3a48a017039919c1448af7adac0596dbade8c23d9e9b20ed

Observation a80e9000-c40a-4121-964b-61f1b9dc0c67 · outbound

This paper cites Transforming and Combining Rewards for Aligning Large Language Models.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.712536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.712536Z digest=sha256:339637155cc8e8db1cc5a1a556ecb55f797fd9cf0ba6ffc70dee2addb4fac793

Observation a26d0694-f2fd-4af0-b501-231962d711fe · outbound

This paper cites Ethical and social risks of harm from Language Models.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Ethical and social risks of harm from Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.716752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.716752Z digest=sha256:efd1ea79bcbc744326a6077a7f18ac28b707b5e80e0057ab0d19c9b4c516f0ce

Observation 9c7c0d23-7c6c-47ca-9aef-7828554af0ff · outbound

This paper cites an unresolved cited work.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.721017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.721017Z digest=sha256:8cad4542ceb79ac18765ad941439114ef9690742f304b8c66fb52a634cd25c1d

Observation 146e3566-1b63-4247-bd29-bb6993028f04 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models Fine-Grained Human Feedback Gives Better Rewards for Language Model Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.725087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.725087Z digest=sha256:3c1e12db388d8c63f0f4a7786744c2e7fb39047edfecdbeee4646a6614437cde

Observation 7a496f50-7965-4c1a-b2dc-cddd7aa76bc5 · outbound

This paper cites online" 'onlinestring :=.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models online" 'onlinestring :=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.729401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.729401Z digest=sha256:f8ff9408179744d46b98391dcc8a9a61b39b71ff022d07d3dd01197a0c9497f4

Observation 6b515aae-a9cf-4cbc-919a-138da253c004 · outbound

This paper cites write newline.

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models write newline

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:45.734231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:31:45.734231Z digest=sha256:2bc965db8bb6b51ce92fd4573c8a5d9a762f4da9e03c9affeb987554fd9796a6

Pith citing papers

No inbound Pith citation observations are available.