Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:31.776781Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.02553.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:31.776781Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:30.092931Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:40:41.514103Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a1610d62-8496-4528-b307-6534336b9c74 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective GPT-4o System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073987cd-41cf-4941-a263-ca7b0bcbdc3c · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Gemini: A Family of Highly Capable Multimodal Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d847781b-9233-4555-8232-c1e126abbad6 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Qwen2.5 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac76498-5a66-438f-8048-394162622066 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective The amazon nova family of models: Technical report and model card
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cfe9e74-aa5e-4339-a953-161932231c2a · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective OpenAI o1 System Card
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 838fd653-fbef-4b82-b5c1-7cc2b920e801 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206e65fa-4c05-4e23-83ff-e6b70f85acf8 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a22d8ee-ed50-44dc-9392-7f7f9866d0bb · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021, 2020
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db41d348-05c2-4c17-b586-f1c328306197 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa568fdd-33ba-47a7-8f3a-c6a07f2bdd31 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d7bb58-110a-4dc1-b1a8-d6fdf26249d2 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac4235ba-1686-4383-89d7-670ef940fb6a · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c31fd329-18ae-4c4a-af35-5ef3063ac485 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d4739b4-8fbb-4f16-9d85-ca8afd503afd · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc7d6f1-aa41-45d1-a570-87e7f284de1d · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Process Reinforcement through Implicit Rewards
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e21099-45e5-456f-9a22-a89f33de763d · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 226e8b7d-71ea-4786-aa9b-ea8cb8881666 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f06674a-aab8-4c2f-ace5-3d413506452c · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Reinforcement learning: An introduction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3eae744-e475-416a-b86c-3b11a18d485e · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Analysis of on-policy policy gradient methods under the distribution mismatch.arXiv preprint arXiv:2503.22244, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebc8704f-7ad7-42cb-b105-420a41cba4df · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective High-dimensional continu- ous control using generalized advantage estimation, 2018
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 229e0672-ac94-41c0-aad3-7d737231e735 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faea92d1-0529-408c-8cf3-62950293087a · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Gonzalez, Hao Zhang, and Ion Stoica
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d855c2e-f274-4eab-97d1-d23f6041ab13 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d37a5d8-7487-4031-94a2-c63c0b08b245 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06c03cc1-5efd-4109-ad8b-30f0011fe3dc · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Multi-turn Reinforcement Learning from Preference Human Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef9d1d0-0259-4636-8775-80a7ff65f737 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04a9702-4b69-44c0-ac3b-86915a75c762 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c7a99b-9411-4c48-9a88-2d440dd9bf5d · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d316c10c-5751-4cb2-a82f-c5351ec05453 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c96c6b6-cb23-4153-9448-1ad05c0b0ae9 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Interactive evaluation for medical LLMs via task- oriented dialogue system
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ed50a14-fea9-4537-9fe4-15c531531cc2 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Let’s verify step by step, 2023
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0a2145-8133-4678-ac20-be71fc5c50ae · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Unraveling rlhf and its variants: Progress and practical engineering insights
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4d243be-85b3-4deb-8dd7-850e0b14cfb7 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Solving math word problems with process- and outcome-based feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce13e47a-3bf7-42d3-bbf8-efa2c0227fb7 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e6e972-8ff4-4c20-a920-2dfcddb4ecf9 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70da07b1-0cf1-44f1-ba5c-0c3613c1b6ab · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005c73bf-7ead-4a34-a89f-163671031cb4 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867e3e6f-9b30-47c1-b77b-cd076a11bcef · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb286a80-c534-4ce2-98f5-35563562af88 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Chatbot arena: An open platform for evaluating llms by human preference
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71bc59d4-7290-4a87-b902-305f8be53e54 · outbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458b6a9e-4f12-408c-853b-786e6f2ba5ed · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
Reference 252
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cff38d40-0384-4dfd-b50f-4b7a1be48da5 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389afa80-1000-430a-9628-73a046081c5c · inbound
Relative Score Policy Optimization for Diffusion Language Models Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.