Pith. sign in

Paper Citation Record · LEDGER

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.02553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02553 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:31.776781Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:30.092931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.514103Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1610d62-8496-4528-b307-6534336b9c74 · outbound

This paper cites GPT-4o System Card.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective GPT-4o System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.119855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.119855Z digest=sha256:a91946c5bfd8baec17e1062cf515a5533c78b2cc3f5597a7fb36cb47ca87a9d8

Observation 073987cd-41cf-4941-a263-ca7b0bcbdc3c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.178328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.178328Z digest=sha256:09752fbb20a13576bc75b1f293328b3205127316cf0e5932cbb0aab1a6bf44a9

Observation d847781b-9233-4555-8232-c1e126abbad6 · outbound

This paper cites Qwen2.5 Technical Report.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.272952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.272952Z digest=sha256:1acfbbc1fbac1de884ff1c0e7adaaf0fd13932dab03f0cd3ba5b6591f309f0bd

Observation fac76498-5a66-438f-8048-394162622066 · outbound

This paper cites The amazon nova family of models: Technical report and model card.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective The amazon nova family of models: Technical report and model card

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:33.438977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:28.391741Z digest=sha256:b9cd82e704409c7c76c25aaefd344e56e0ff8cb0d10c3ba32edae161913e08ac

Observation 8cfe9e74-aa5e-4339-a953-161932231c2a · outbound

This paper cites OpenAI o1 System Card.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective OpenAI o1 System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.487274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.487274Z digest=sha256:1bd45a1ed8c53f47bbe2884c419632b776c8db12a8eda05b6c1dbfa991d1c8c8

Observation 838fd653-fbef-4b82-b5c1-7cc2b920e801 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.616012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.616012Z digest=sha256:317c548f753b6c1ac2a0b1e200f5feda66f473c994cac42d753caf36c121f963

Observation 206e65fa-4c05-4e23-83ff-e6b70f85acf8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Proximal Policy Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.675759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.675759Z digest=sha256:97ea72a5913e643d1a9ae7040ab17b2019da603aa708430e3d781ee938de3f86

Observation 7a22d8ee-ed50-44dc-9392-7f7f9866d0bb · outbound

This paper cites Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021, 2020.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021, 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.759565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.759565Z digest=sha256:985c8ada4a48bfecbfe14836a5704bb944749f4c24551583ae068f7aaaa9e277

Observation db41d348-05c2-4c17-b586-f1c328306197 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.849369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.849369Z digest=sha256:dbef59b4226007c48cdb038823cb299fdbb2f4080c00f65d2d9ccfee31726060

Observation fa568fdd-33ba-47a7-8f3a-c6a07f2bdd31 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.936948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.936948Z digest=sha256:71ad0a6a606d4ca0dfb53ad697a96cc32b92ef8d7226259f8e24c295afb7610b

Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.011230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.011230Z digest=sha256:5fae6af789e08fd7692a2f19cf5f3f98b87aff4feb08961e373c4110dade651d

Observation 92d7bb58-110a-4dc1-b1a8-d6fdf26249d2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.157791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.157791Z digest=sha256:e5b85680465e27f0e287019e427fd74f57a2e3bd710307e9ceaa988dfaf08483

Observation ac4235ba-1686-4383-89d7-670ef940fb6a · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.285863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.285863Z digest=sha256:b3067966a98790c0ca71cecf6c73f4d46dd3a63d1b33b36205954b7fc4c284e3

Observation c31fd329-18ae-4c4a-af35-5ef3063ac485 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.379167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.379167Z digest=sha256:613da14e69e611420ca478380fdbedd05e1e9fbfbcca136db06c322ebba17d02

Observation 3d4739b4-8fbb-4f16-9d85-ca8afd503afd · outbound

This paper cites Dense Reward for Free in Reinforcement Learning from Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Dense Reward for Free in Reinforcement Learning from Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.474618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.474618Z digest=sha256:18a3a6c57084ea9f88e5518d9ebbb74738b69276206d7e6a5d7e4ed575f03e63

Observation 9dc7d6f1-aa41-45d1-a570-87e7f284de1d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Process Reinforcement through Implicit Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.517440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.517440Z digest=sha256:72e9cbf923aa02b4e06cf0bce3cae7aef1aebd7dfa134adca3b72ed22380420e

Observation d1e21099-45e5-456f-9a22-a89f33de763d · outbound

This paper cites RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:32:32.289654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:29.571593Z digest=sha256:5d6310d791db167a7ae034d2c610525bd27808d42a5d9f50649e2129d2e758b2

Observation 226e8b7d-71ea-4786-aa9b-ea8cb8881666 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.652497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.652497Z digest=sha256:5068559f98dd88756235c5b77ee40e95f3483c0ae19f728807921eedbaa51475

Observation 0f06674a-aab8-4c2f-ace5-3d413506452c · outbound

This paper cites Reinforcement learning: An introduction.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Reinforcement learning: An introduction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:33.106505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:29.736929Z digest=sha256:d6d6512b297cb654adcad6e6c2339cb475d8b853afe5d1fbfeb771278a4fb5a1

Observation e3eae744-e475-416a-b86c-3b11a18d485e · outbound

This paper cites Analysis of on-policy policy gradient methods under the distribution mismatch.arXiv preprint arXiv:2503.22244, 2025.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Analysis of on-policy policy gradient methods under the distribution mismatch.arXiv preprint arXiv:2503.22244, 2025

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-07T11:32:32.175493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:29.838106Z digest=sha256:0882900e903d4607965d9a75954eed6ff8da6645ac4298f6cbbda33fdefd51ee

Observation ebc8704f-7ad7-42cb-b105-420a41cba4df · outbound

This paper cites High-dimensional continu- ous control using generalized advantage estimation, 2018.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective High-dimensional continu- ous control using generalized advantage estimation, 2018

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.986611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:29.935825Z digest=sha256:974aea2df63e7cb4fcdb98a237e64c39091a6f912baee2222f3072a3591f73dc

Observation 229e0672-ac94-41c0-aad3-7d737231e735 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.026404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.026404Z digest=sha256:2dfa47e503a5bdb385cb6d46dcea5ae78a3df66922d719c60a6bcca53e16cb29

Observation faea92d1-0529-408c-8cf3-62950293087a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Gonzalez, Hao Zhang, and Ion Stoica

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.149593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.149593Z digest=sha256:784c8b682e1429805c0269f0875fb36cf8b291a9778fbd85bb1e0338851c9361

Observation 8d855c2e-f274-4eab-97d1-d23f6041ab13 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.869318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:30.236850Z digest=sha256:1b1a6fda8ad4285bd882d117601c5799ea4757aad3198a89db2a862db9437658

Observation 3d37a5d8-7487-4031-94a2-c63c0b08b245 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.314881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.314881Z digest=sha256:04e71b9ca2b0188e6de3663d346c771c8ed75d3446001915a3b0bd5c6d5d34ab

Observation 06c03cc1-5efd-4109-ad8b-30f0011fe3dc · outbound

This paper cites Multi-turn Reinforcement Learning from Preference Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.406644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.406644Z digest=sha256:7d2e043be8bf76e92f03f78ae56bbcb7c3d9c7e4163c4ea20b0594b76999a2d3

Observation 4ef9d1d0-0259-4636-8775-80a7ff65f737 · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.490860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.490860Z digest=sha256:d32f231fcf4d7dbc2a96de474bafd739f8357405ee2ef168d5b581d6b199c047

Observation f04a9702-4b69-44c0-ac3b-86915a75c762 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.549942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.549942Z digest=sha256:d5fd68c951e59341a15cea0f1c867747d18f4dbc3ecc9a17f7ca0d60eca25731

Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.615635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.615635Z digest=sha256:0cca8c04f15176c864e19eba635cac36a4e16982d228caac163159569b4574ac

Observation 46c7a99b-9411-4c48-9a88-2d440dd9bf5d · outbound

This paper cites PlanGenLLMs: A Modern Survey of LLM Planning Capabilities.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.682383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.682383Z digest=sha256:2614a3aaf3f476434624bdb66dbd93032421e7560ebe921b387a84dd1f7d9813

Observation d316c10c-5751-4cb2-a82f-c5351ec05453 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.761569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.761569Z digest=sha256:6d6f57c9f9ee72d47faf99d5293faf3c9e897cf8516b12de68cae0d91af47b13

Observation 4c96c6b6-cb23-4153-9448-1ad05c0b0ae9 · outbound

This paper cites Interactive evaluation for medical LLMs via task- oriented dialogue system.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Interactive evaluation for medical LLMs via task- oriented dialogue system

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.721910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:30.841930Z digest=sha256:26f4752438b4fdd2f758c032d09b17817042b0b76c0a7504b47c7437d3a7a85c

Observation 7ed50a14-fea9-4537-9fe4-15c531531cc2 · outbound

This paper cites Let’s verify step by step, 2023.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Let’s verify step by step, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.927433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.927433Z digest=sha256:2f27dc29eb1e9df5ed54e8d6b5c078f9e2a4f2241c234db19f03fec5c4d838e8

Observation 7a0a2145-8133-4678-ac20-be71fc5c50ae · outbound

This paper cites Unraveling rlhf and its variants: Progress and practical engineering insights.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Unraveling rlhf and its variants: Progress and practical engineering insights

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.554729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:32:31.056825Z digest=sha256:556f367317deff3964e143adc6447643ac776db54d7e89724a59732270de64eb

Observation c4d243be-85b3-4deb-8dd7-850e0b14cfb7 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.151918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.151918Z digest=sha256:58b32c207ce5c2743e6412b4c44d7e5e3ffec95828b777ad61177802bd1e78d0

Observation ce13e47a-3bf7-42d3-bbf8-efa2c0227fb7 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.223415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.223415Z digest=sha256:f2d91a2db5c4da9bd76fdd3f9a4f850c694f5a7a81d2a682c7337c836db98485

Observation 86e6e972-8ff4-4c20-a920-2dfcddb4ecf9 · outbound

This paper cites Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.310922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.310922Z digest=sha256:883347bc7b770eb68219c0624544fc6142223435a98fe971280f1dda6d295046

Observation 70da07b1-0cf1-44f1-ba5c-0c3613c1b6ab · outbound

This paper cites TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.365581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.365581Z digest=sha256:774c5e704fb34c222c23996da69e40c32d0e4c23994c0647a27656824b6419a7

Observation 005c73bf-7ead-4a34-a89f-163671031cb4 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.431529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.431529Z digest=sha256:ab814849da9f2c783a7345444d7a242e81006e9a87743e99681e134249d686ab

Observation 867e3e6f-9b30-47c1-b77b-cd076a11bcef · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.527587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.527587Z digest=sha256:c7524d1e29183e7e18039c0bb7b1baf7cac044593bc6dc7cc90044310959911e

Observation eb286a80-c534-4ce2-98f5-35563562af88 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Chatbot arena: An open platform for evaluating llms by human preference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.601560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.601560Z digest=sha256:7228397fa015e445594002670c0f4524c76ba1a6242e2a983fb8cea3a423e1f6

Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.693251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.693251Z digest=sha256:1cdf0544b41ffd7c76a04e7ae766102af14109f7b6e25ce168a22fa6ca030052

Observation 71bc59d4-7290-4a87-b902-305f8be53e54 · outbound

This paper cites MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.776781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.776781Z digest=sha256:b8cb60f0db44697507ee66eb036c95aa28786296435efe017dc007adbb66d19a

Pith citing papers

Observation 458b6a9e-4f12-408c-853b-786e6f2ba5ed · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 252

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.517098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:4dba68995edf6c73411d90b2537042cc42d2786613e7744247c8af5cdaeceb99

Observation cff38d40-0384-4dfd-b50f-4b7a1be48da5 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.092931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.092931Z digest=sha256:cacb63b0353c3ea50dfcd326a91f16bd27a73e8d80b62ca1ec9bb4830aff27ec

Observation 389afa80-1000-430a-9628-73a046081c5c · inbound

Relative Score Policy Optimization for Diffusion Language Models cites this paper.

Relative Score Policy Optimization for Diffusion Language Models Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:29.690474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:47:42.196931Z digest=sha256:7f0481dbb466ff06198de7031eac3efacb2e555bfe2a973d1565bf953f2e62e5