Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:50:35.003281Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2411.08302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:50:35.003281Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:29.571593Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:32:32.239050Z
69 of 69 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3e6acc89-d600-4382-ad3f-53783c4f1d87 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529bdd81-f1b2-4174-8074-5d25afcf42cf · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6beb3645-bde9-4a55-a7ad-9c1001fb72af · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52c0f83d-ea05-451b-8564-8a10fb410a38 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 12c3585a-d123-417a-bf17-50032316d345 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f047105-0b74-4c26-b34f-2018c66ab707 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e998b9-60a2-473d-a47d-e279dda7c4d4 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c1b531-b003-4f13-ac77-13b13ffd6c28 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 62e52f70-9975-4f2e-9d68-99ea9e292bfa · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 741bf821-0629-4dbc-a4aa-84efaf768afc · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19682fda-da04-4d1c-b6b6-e8a68c98d55e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 743ee7bb-c825-48fb-bf95-b3067261dc35 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b19291-0627-4ce5-aed2-06ead1ac4319 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f4d930-edd4-4138-a814-55e6ce53954f · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33252586-03b0-4b13-ae9c-121fe6795f49 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution KTO: Model Alignment as Prospect Theoretic Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d47e865b-a0af-4f9c-b76c-4f2a97402efd · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68c2767b-1198-4456-865f-7be8311d6ee2 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9d3fccc-4f2c-41c7-a872-ccab051798b0 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526f53eb-5e4b-4139-b0ae-308906e99bb7 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Aligning Language Models with Preferences through f-divergence Minimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b068cb7c-86c9-44e3-b016-f01eb40a6b50 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6cc7319-615d-4a20-8927-d3ab4264d56e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution ORPO: Monolithic Preference Optimization without Reference Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027b6d57-d2fe-414a-90a1-03ebb53350df · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c492ae-3d41-43d4-b6fb-628122d4dbd0 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c71c181c-cb83-455e-b5ba-95452d2dd40d · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution u chemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan G \
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 898702c6-e737-4651-8b91-939f6f63ebeb · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution ARGS: Alignment as Reward-Guided Search
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822157e9-8fe2-4d20-9d82-253c43f63f2d · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85763b68-e586-4bc5-a7cf-73babd44f39d · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 475c6d96-87bc-43fd-9cdc-3909a0cf2eff · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 93c1f63e-6e9b-4f8c-8ac4-9da52ba364f8 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8624b25-96e5-489f-83f2-c3df85dc8c15 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a7c470e-670b-4ffe-8849-e99afec5af0e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e16f94f-412a-48cc-b8d7-644e20b72a39 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eb6c0c4-867f-45e5-a7fb-311d3ae5bcbf · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c824ec30-1705-4773-92cb-e167098feb8e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Statistical Rejection Sampling Improves Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74cff5f9-ca3f-4c61-8062-f07a0706a2c5 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22771146-9716-4e82-99a5-eb2e0f3b1d5b · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2222bff-4d8a-4790-a461-d070853b2a23 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55079465-cbe2-4c50-911e-8739e9f44d4e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ae341bb-bcf6-42a6-9479-b3c3df5902cf · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a5c017-9d59-43b5-8171-b5fe7fbb0687 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b340e3-5408-475c-92a1-d6836787694b · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Disentangling Length from Quality in Direct Preference Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aedc0e4-1165-4e9b-8abe-52c6a9b5fada · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab7e6fa-1638-49f0-94d0-4b074aa82a05 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60b61f3-232e-4d21-9acf-09bfbd5d3421 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af80aae-f344-4124-9eb5-3f2825065f94 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822c7bc9-7312-4640-9dd0-17df0ca83c38 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d17107f-8e5b-4b2e-9bac-96897211ffb8 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Proximal Policy Optimization Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10fbcac8-07ec-493b-b48f-396e816554d2 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3631417-6f14-496a-91ae-5508eac68a06 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2612349b-3471-4c34-ba10-93683569a6f9 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae413a4-83c0-4c83-b02c-43bfe1c2ad95 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Qwen2.5 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d476f158-ab73-42e8-875b-331474b937b8 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution LLaMA: Open and Efficient Foundation Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 161ed8c9-fb14-4521-89a8-28bd4f5734ad · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2601f6b0-13c2-42e8-be1f-fbbc0f7b6877 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444964d3-ef52-4455-ac4c-984598978250 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41dbd110-1378-4f96-9557-731e9b4cd06b · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2cd10a1a-8dcc-483c-aaed-c7542a4dec09 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8953b6-fc9a-4fb7-84d8-966ebc882a67 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0129bce-d5ba-49c3-95c8-b328904045f5 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d278867b-dce8-4b71-81ec-e8e1c06a7c5d · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375dbfc0-6468-48a1-a956-ca2b058a13e9 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Benchmarking Machine Translation with Cultural Awareness
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed21aca3-f162-4302-a4ed-978bbaa9d021 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3b4dc6-250b-4510-9daf-1db2eaaaed97 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 80e4268d-f709-4dcd-893c-ce3f06ca8c0e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66dc5680-099b-439a-8e2a-4c969731c125 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21877702-e9f5-4e27-b7bb-c48cc6b28c95 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202df54d-24fa-46e0-86db-969ff2b89260 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Fine-Tuning Language Models from Human Preferences
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed266158-72da-47c1-82cc-701de3d149d3 · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution online" 'onlinestring :=
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186ae0d9-144c-4bc5-8f65-250065a4c66e · outbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution write newline
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e21099-45e5-456f-9a22-a89f33de763d · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b6b31347-93a8-45a4-91e0-514061b8e5b6 · inbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.