Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:18.716371Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2505.17218.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:18.716371Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:06:50.901959Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T08:04:29.012591Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f2fa5f58-4f1f-4daf-b021-5e534733bb6f · outbound
Effective Reinforcement Learning for Reasoning in Language Models online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97841ca2-85be-437f-b2a4-1f9298a55adc · outbound
Effective Reinforcement Learning for Reasoning in Language Models write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0692bc-14ce-48d8-b019-878a3bfab9f4 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301d928e-21e4-4ad1-be3a-a639961f7515 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81179bea-7776-4a74-bbd2-e377439404e0 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a52da57-1d21-4abd-889f-50538f316495 · outbound
Effective Reinforcement Learning for Reasoning in Language Models FireAct: Toward Language Agent Fine-tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7537b4e-1f81-47a5-bbc7-9a03136413ba · outbound
Effective Reinforcement Learning for Reasoning in Language Models Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39563835-0abf-4359-9613-2f588b7937b3 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f966a498-ec37-4cf1-9e01-3eccea43e577 · outbound
Effective Reinforcement Learning for Reasoning in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1df620-03e1-4f3e-8d84-d1bd9aef3e45 · outbound
Effective Reinforcement Learning for Reasoning in Language Models DeepSeek-V3 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f4e25b-0026-4598-9290-4d4737d2310e · outbound
Effective Reinforcement Learning for Reasoning in Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050c12db-bcab-42ea-804c-9a2547a3080e · outbound
Effective Reinforcement Learning for Reasoning in Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eccfac8-72d1-48ae-a614-79ee08b6f4d0 · outbound
Effective Reinforcement Learning for Reasoning in Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ce54ae-9fc4-43f3-b569-e7d0796fb5b8 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b5de8a-0e45-4078-b03c-0c1fcbee1a0e · outbound
Effective Reinforcement Learning for Reasoning in Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b12b335-90f4-4257-9d3f-91771ac5fa9d · outbound
Effective Reinforcement Learning for Reasoning in Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 733bbca4-b1e2-4ed9-8c6a-204e78ee2a8f · outbound
Effective Reinforcement Learning for Reasoning in Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e517c5bc-b108-4c77-9ac9-9608e45ca82c · outbound
Effective Reinforcement Learning for Reasoning in Language Models Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 270e01a2-50f5-47df-a1bd-76d15e791718 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Let's Verify Step by Step
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042a3e37-ed0d-4921-8777-21f179c4531d · outbound
Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e601eb6-26b0-4b54-bddb-8822e369e2dd · outbound
Effective Reinforcement Learning for Reasoning in Language Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814140a9-8ba0-4f56-bf05-c33cc0f80a66 · outbound
Effective Reinforcement Learning for Reasoning in Language Models s1: Simple test-time scaling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43182d4f-8861-4cee-83b8-1500f8c0036a · outbound
Effective Reinforcement Learning for Reasoning in Language Models Self-Imitation Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3f68f129-c728-4c2e-a650-2d86c599d465 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545294a5-5926-460a-b977-543dba9b152a · outbound
Effective Reinforcement Learning for Reasoning in Language Models CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780eb4ef-b32b-40c1-95a6-d78ffba76438 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Qwen2.5 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e56e107-4a44-4340-ba18-a94071a5042e · outbound
Effective Reinforcement Learning for Reasoning in Language Models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639f44f9-78aa-41e0-aa57-6ad4818c4c78 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Trust Region Policy Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfa1d65-2157-41e4-9722-60757a500681 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Proximal Policy Optimization Algorithms
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac161ae9-b028-48b7-a580-d042a309d5dd · outbound
Effective Reinforcement Learning for Reasoning in Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a691dc3-28f2-4255-be38-cd450d3a1558 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbbb8d0-7fd2-4f2a-b8aa-45c296fd114c · outbound
Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2659db-8d4f-42df-ae76-e81e7824c3df · outbound
Effective Reinforcement Learning for Reasoning in Language Models Sutton and Andrew G
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da98b5e-23a4-40ed-b66b-21c67f6ce3bd · outbound
Effective Reinforcement Learning for Reasoning in Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fbd850b-24ad-4320-b377-21e8ef4a88a9 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16523a4-1438-4de7-833b-e12e8e102578 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe47ffdc-66f5-4879-8e9c-6fde23034d61 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Base Models Beat Aligned Models at Randomness and Creativity
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd47ed01-d16f-429a-a201-48e1db98f6d8 · outbound
Effective Reinforcement Learning for Reasoning in Language Models Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c10c2e-0fc2-438c-8532-43fa80c9e67c · outbound
Effective Reinforcement Learning for Reasoning in Language Models ReAct: Synergizing Reasoning and Acting in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eca98ce-8d50-4895-9a5f-eb59a07aa06b · outbound
Effective Reinforcement Learning for Reasoning in Language Models AgentTuning: Enabling Generalized Agent Abilities for LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25824708-5506-4a7e-bee0-b07e9bc130f3 · outbound
Effective Reinforcement Learning for Reasoning in Language Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994ae51c-3fec-4000-86d4-ddd81495384d · inbound
PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF Effective Reinforcement Learning for Reasoning in Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c7e51b63-acdb-44de-a9fd-2741712cc433 · inbound
SLPO: Scaling Latent Reasoning via a Surrogate Policy Effective Reinforcement Learning for Reasoning in Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.