Pith. sign in

Paper Citation Record · LEDGER

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2411.08302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08302 v2

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:50:35.003281Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:29.571593Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:32:32.239050Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e6acc89-d600-4382-ad3f-53783c4f1d87 · outbound

This paper cites GPT-4 Technical Report.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.474159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.474159Z digest=sha256:8543e5601edf285268fad566e1a55e5325bdb3fcbeec019bed923a0a131b4a51

Observation 529bdd81-f1b2-4174-8074-5d25afcf42cf · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.552488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.482033Z digest=sha256:87e5014d92dd9fa19640e39515c01fa3ade9f25cfed755c07f470c847fbfa250

Observation 6beb3645-bde9-4a55-a7ad-9c1001fb72af · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.488878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.488878Z digest=sha256:9ee8f450ddab7bb9632cbcf6f9b0a83ac7a91726fa4fba7dc57a3463d1bcf618

Observation 52c0f83d-ea05-451b-8564-8a10fb410a38 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.519898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.494804Z digest=sha256:f3ef8734e8354ead73eee27d9d711e9cd2c8534dcc91d74b1890bcc90971c426

Observation 12c3585a-d123-417a-bf17-50032316d345 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.500748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.500748Z digest=sha256:e89f4bc4172bf25479f7d5bd752eafae486a1bec2da72549809f078ccdb0cd2a

Observation 1f047105-0b74-4c26-b34f-2018c66ab707 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.507166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.507166Z digest=sha256:b8e79f130bda49e2ad59b72e9157d555abc91611ca18e163a9882ca9f0fa2977

Observation 16e998b9-60a2-473d-a47d-e279dda7c4d4 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.513493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.513493Z digest=sha256:059c822e54d4f5ce609f6f325ecc8d004c65c00157fa8d6267795cd9a2cf0a7e

Observation 42c1b531-b003-4f13-ac77-13b13ffd6c28 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.465093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.520534Z digest=sha256:1b7b720d91f7af97d36c986deb9187a2fb6c2c606da100b03bce3f9e155edb82

Observation 62e52f70-9975-4f2e-9d68-99ea9e292bfa · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.444823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.528750Z digest=sha256:0cee8f10e7a38d6706b386f7504cd21dc87dfd5d77ae9a6c605af60e51b7801a

Observation 741bf821-0629-4dbc-a4aa-84efaf768afc · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.537484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.537484Z digest=sha256:1300cb539f5314f2ed6a81ae6d739c240341f76b0262fd7eb2ed7785e0a66c27

Observation 19682fda-da04-4d1c-b6b6-e8a68c98d55e · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.410882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.544450Z digest=sha256:1253416cd5e476e0f52ca4823cbabae16e7eb151eee8491de053ffbd2a09ba8b

Observation 743ee7bb-c825-48fb-bf95-b3067261dc35 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.550946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.550946Z digest=sha256:df242fb8f07c6afe1323d9f06a633308da521ba926bc56e74bdf263eeafdc209

Observation 27b19291-0627-4ce5-aed2-06ead1ac4319 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.560553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.560553Z digest=sha256:e908dee4ff183c511d205f066e47dd6e285d6eab76381133c58fc9c6df1b7cf0

Observation 37f4d930-edd4-4138-a814-55e6ce53954f · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.359271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.568786Z digest=sha256:ab78355c6f4ccc224f4713966f3ba08ca7e616c5696fd7bc496c2b233ee844ae

Observation 33252586-03b0-4b13-ae9c-121fe6795f49 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution KTO: Model Alignment as Prospect Theoretic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.576637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.576637Z digest=sha256:55d69b12a37dea4c3e9d681acc122dc4a4e99932cbb526510bc2dd15968c2b42

Observation d47e865b-a0af-4f9c-b76c-4f2a97402efd · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.339757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.586729Z digest=sha256:639fea6cde6697580c96e9e194717e52564f7da18383e79786973712cf1be727

Observation 68c2767b-1198-4456-865f-7be8311d6ee2 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.320152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.592392Z digest=sha256:372995591c6124d5a7ddd5dd5749fcbadd80b41c6a32975301e05e9a03ab6eae

Observation d9d3fccc-4f2c-41c7-a872-ccab051798b0 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.597954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.597954Z digest=sha256:922a45cc4d0bcd461e354ce505df7bbc0b4e00db78b25836135f26a8bc0092ef

Observation 526f53eb-5e4b-4139-b0ae-308906e99bb7 · outbound

This paper cites Aligning Language Models with Preferences through f-divergence Minimization.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Aligning Language Models with Preferences through f-divergence Minimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.606959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.606959Z digest=sha256:0d4b5f244c7fd62811a420cb9edf5ea23df34f6e396ac697602f2bc77e4f296c

Observation b068cb7c-86c9-44e3-b016-f01eb40a6b50 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.613515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.613515Z digest=sha256:79f1658827126575533a2d9d1ec5870e2502e8bd2121b091d237b3cef27d0df0

Observation e6cc7319-615d-4a20-8927-d3ab4264d56e · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution ORPO: Monolithic Preference Optimization without Reference Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.622974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.622974Z digest=sha256:72651465232b55ee56678223919d92569801d3e2303db132cfaba6e4e29972bb

Observation 027b6d57-d2fe-414a-90a1-03ebb53350df · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.629915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.629915Z digest=sha256:a3be058b2fb6bbc67bf0c2e8b7dfefda129d0ca07ddb81f3bd942905c52d18dc

Observation 19c492ae-3d41-43d4-b6fb-628122d4dbd0 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.297658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.638941Z digest=sha256:20dba74eb3ede850644acafe20edaff317452fda29d7e593dd07273cfcabe8f8

Observation c71c181c-cb83-455e-b5ba-95452d2dd40d · outbound

This paper cites u chemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan G \.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution u chemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan G \

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.644998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.644998Z digest=sha256:c99ec717fb521bbb40e0d91bd58a341b14a0725b5f20442f7e88d4e5f901edbc

Observation 898702c6-e737-4651-8b91-939f6f63ebeb · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution ARGS: Alignment as Reward-Guided Search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.653892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.653892Z digest=sha256:7a2ce22b09b4cff6061e648c923788b1ac2b591283aff0c847b65fd0718033f2

Observation 822157e9-8fe2-4d20-9d82-253c43f63f2d · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.662773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.662773Z digest=sha256:0e2992e28873db909e8412333844471a98de6c6bd80d0dec04df1f3d0af0687d

Observation 85763b68-e586-4bc5-a7cf-73babd44f39d · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.249540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.670548Z digest=sha256:06351a8ecf6ea84734bb3d848c9c8ace495b583ac2086c6f41b3b21c0fb68725

Observation 475c6d96-87bc-43fd-9cdc-3909a0cf2eff · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.219705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.681796Z digest=sha256:f016d5d8882e9248ba93ed435d909888240da8864ce640a0aba56269bea903c3

Observation 93c1f63e-6e9b-4f8c-8ac4-9da52ba364f8 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.696444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.696444Z digest=sha256:708e50495034498962966cb450d66f3a2fd729be4970ae1d7f24cff58abd5b0a

Observation e8624b25-96e5-489f-83f2-c3df85dc8c15 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.179992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.702797Z digest=sha256:f921ca71ee77198ebb172d8e40663282ab55f01db5ff3a6fef0c949fd8868c1e

Observation 5a7c470e-670b-4ffe-8849-e99afec5af0e · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.708499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.708499Z digest=sha256:eb47b1bdaec4d1005ed89c7d7ed7133e4d7a2cd47b76df7b2119a5635546abd6

Observation 7e16f94f-412a-48cc-b8d7-644e20b72a39 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.714984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.714984Z digest=sha256:e0906b57e05ee87386dcb782e1acb2caf76f692b48c6cd5d8513a86c08142b1a

Observation 1eb6c0c4-867f-45e5-a7fb-311d3ae5bcbf · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.721768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.721768Z digest=sha256:fbbf3f05ec0666a81664118d6f11c521b9cea768603c0ab8ce9d5779c266166b

Observation c824ec30-1705-4773-92cb-e167098feb8e · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Statistical Rejection Sampling Improves Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.729198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.729198Z digest=sha256:28efcc22a9fb80b52c6afc9e53c25374eab6a25d5dac2f97492ad179698f87c3

Observation 74cff5f9-ca3f-4c61-8062-f07a0706a2c5 · outbound

This paper cites Extensive Self-Contrast Enables Feedback-Free Language Model Alignment.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Extensive Self-Contrast Enables Feedback-Free Language Model Alignment

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-12T21:50:35.510693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.743396Z digest=sha256:d46c83fed7056e1d34f709e46a6b107baba71d2dd77acc875f5095cca2c64925

Observation 22771146-9716-4e82-99a5-eb2e0f3b1d5b · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.757342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.757342Z digest=sha256:5a866d42e2b64f1b3fc35336e21da98e28293699c17e7aa3fc6aabf6e2e7296c

Observation d2222bff-4d8a-4790-a461-d070853b2a23 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.124671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.766572Z digest=sha256:899258a20c839500922c29e530f420c751bdb99b5a706c1028a314d55ffa0b5d

Observation 55079465-cbe2-4c50-911e-8739e9f44d4e · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:36.103862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.772603Z digest=sha256:8f5ef8f1582e17cc496734bb45fa29a2ee8a97013ff7738b97771659da62d91d

Observation 8ae341bb-bcf6-42a6-9479-b3c3df5902cf · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.782554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.782554Z digest=sha256:6d68f14082b29524224ad35250cf1161d121e2b1df5e6feb811a2c94c67aefc9

Observation b8a5c017-9d59-43b5-8171-b5fe7fbb0687 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.790041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.790041Z digest=sha256:f8b6b4bf96feb9fd0035c767872269669ebd5fff83f69ba716f29c4fb813a2db

Observation 79b340e3-5408-475c-92a1-d6836787694b · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Disentangling Length from Quality in Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.797495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.797495Z digest=sha256:3196942a6fada96acceca2ade8ffd802066e748d62bd9a9874534db53a512adf

Observation 0aedc0e4-1165-4e9b-8abe-52c6a9b5fada · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.805170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.805170Z digest=sha256:ec7f605b63f62e6c8cf3d803114aab7552c211047e1c2825d2c33b9f6ad4adb0

Observation 5ab7e6fa-1638-49f0-94d0-4b074aa82a05 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.812387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.812387Z digest=sha256:e6f6528595ac346a1494a336f780542a8fbfa2bbc5945b994523fd4ac3c4851c

Observation f60b61f3-232e-4d21-9acf-09bfbd5d3421 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.821043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.821043Z digest=sha256:5976fb5ded8ec0d12b1848c76b0d925c2303c5d37295767c2ce575ca9bb46f58

Observation 8af80aae-f344-4124-9eb5-3f2825065f94 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.829523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.829523Z digest=sha256:2e57d620158be482cf509d2404c08a99d5570b52f59565913302ab2306db2de0

Observation 822c7bc9-7312-4640-9dd0-17df0ca83c38 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.837659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.837659Z digest=sha256:e3ea21105386013865ae959524a97360347af4a59fcbcc04f93a5f0d1bc0cf0d

Observation 3d17107f-8e5b-4b2e-9bac-96897211ffb8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Proximal Policy Optimization Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.845174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.845174Z digest=sha256:d98056fd9f1dc21508dafaf934ec4fe9cdf5302921c90847f8bc63d59cea20f6

Observation 10fbcac8-07ec-493b-b48f-396e816554d2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.852425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.852425Z digest=sha256:d49f5e645e3447d49c292cc69b820ffc9d84d62375525e952f84d1432f5c4455

Observation d3631417-6f14-496a-91ae-5508eac68a06 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.862165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.862165Z digest=sha256:0e3f1fd83012bc384c9665511edd29e2af537085495b4028ccd5106153145712

Observation 2612349b-3471-4c34-ba10-93683569a6f9 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.868299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.868299Z digest=sha256:a0f710e00078a25a29899ca618803990e7a953ad9e58294493a20934e9a6b7ef

Observation bae413a4-83c0-4c83-b02c-43bfe1c2ad95 · outbound

This paper cites Qwen2.5 Technical Report.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.875441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.875441Z digest=sha256:6e5229c82e64cac8e5275e53dc2eff921a485b347beed31fab90ec100002b96f

Observation d476f158-ab73-42e8-875b-331474b937b8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.882045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.882045Z digest=sha256:274d71f7f9f569c1d6f85ca35f5ce5ab9201f832140e871d0d0a16b5532add50

Observation 161ed8c9-fb14-4521-89a8-28bd4f5734ad · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.889384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.889384Z digest=sha256:1a7223d5c5526888474f55cbf397c8f0679355aecd2e66ab18b7a8ac47dafa15

Observation 2601f6b0-13c2-42e8-be1f-fbbc0f7b6877 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.896771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.896771Z digest=sha256:0c18d627d9bcbe1307505422b208173951b37974ac9a0235c636880f711c9b35

Observation 444964d3-ef52-4455-ac4c-984598978250 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.903038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.903038Z digest=sha256:1fc7089b46f4f95394d2ad17530810539dcae2ec7cef270b96dfbe896d8ae1e9

Observation 41dbd110-1378-4f96-9557-731e9b4cd06b · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:35.926216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.910326Z digest=sha256:00d78fa20221e37f5275417657dce7cae29c33c9a1acdb444bb5506b96c523b7

Observation 2cd10a1a-8dcc-483c-aaed-c7542a4dec09 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.916325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.916325Z digest=sha256:1d82434f4f749710bfeeb777d752722c7b62bdf51bdcc317bd5cf96100f369b5

Observation 2a8953b6-fc9a-4fb7-84d8-966ebc882a67 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.922202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.922202Z digest=sha256:140faea8113736b6d8fcf955e545463b9e73585442582acd4803e912b8491cd0

Observation b0129bce-d5ba-49c3-95c8-b328904045f5 · outbound

This paper cites Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.927776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.927776Z digest=sha256:d850d789081068ec461dc7dd659deb5eeed11a050079e2b43f2695878f15c534

Observation d278867b-dce8-4b71-81ec-e8e1c06a7c5d · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.933985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.933985Z digest=sha256:7cea9bdc46d10162d334d6d4553a4fd6e249691c869662326afd80af34974fdc

Observation 375dbfc0-6468-48a1-a956-ca2b058a13e9 · outbound

This paper cites Benchmarking Machine Translation with Cultural Awareness.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Benchmarking Machine Translation with Cultural Awareness

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.939837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.939837Z digest=sha256:10559c29883f786b1e497e9777f701bce8b2a98539ab33560606f6ccfe5dafeb

Observation ed21aca3-f162-4302-a4ed-978bbaa9d021 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.946411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.946411Z digest=sha256:25921e0e0e8214f031e97e8294b8096c1e756987f8f40e6f9631e7ab4dcb110d

Observation 4b3b4dc6-250b-4510-9daf-1db2eaaaed97 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:50:35.847476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-12T21:50:34.952716Z digest=sha256:427519fcda605a31ebdae9653deabb8f23e1eaa2da9e3be66471e3a350223c53

Observation 80e4268d-f709-4dcd-893c-ce3f06ca8c0e · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.961970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.961970Z digest=sha256:00146ead1c7da164f02dcfe6cde7edc5053a92b6ff7704aa2815c6671b244fbf

Observation 66dc5680-099b-439a-8e2a-4c969731c125 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.968647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.968647Z digest=sha256:e698b8e8897d2beec9ee1586af1f3c721a76fbdfebfb29c1d8ce5dee7a32f70b

Observation 21877702-e9f5-4e27-b7bb-c48cc6b28c95 · outbound

This paper cites an unresolved cited work.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.977977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.977977Z digest=sha256:1c25f07941405e6454cdacbb62c72a92049dbfba39cc99b8608fa3364e018ed5

Observation 202df54d-24fa-46e0-86db-969ff2b89260 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Fine-Tuning Language Models from Human Preferences

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.987580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.987580Z digest=sha256:d1e03ad54145f64853a6d8b5b612d08c4dd1a0e5a53ec5f66fca996debc4d3cb

Observation ed266158-72da-47c1-82cc-701de3d149d3 · outbound

This paper cites online" 'onlinestring :=.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution online" 'onlinestring :=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.994841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.994841Z digest=sha256:d4eadf08940a03e1b7d4b6cdd365e416d956d432145e2f679ede0e465a44c03a

Observation 186ae0d9-144c-4bc5-8f65-250065a4c66e · outbound

This paper cites write newline.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution write newline

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:35.003281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:35.003281Z digest=sha256:868cc890ef40b8bfd14ac043c2839b62bcf2df8a84f019b70805fca6bbfb75ae

Pith citing papers

Observation d1e21099-45e5-456f-9a22-a89f33de763d · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:32:32.289654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:32:29.571593Z digest=sha256:9845867d9fa77feb7442b4cfe75e8d0035a2a7ba2da50e11eedad753467c3d1b

Observation b6b31347-93a8-45a4-91e0-514061b8e5b6 · inbound

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF cites this paper.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:43.513582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:43.513582Z digest=sha256:0f978bf06017283f173e9e802161fdd2159ae85e2ae3b03938b3ab1b4f553d90