Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T15:49:44.263123Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 4 inbound Pith citation observations for arXiv:2505.07527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T15:49:44.263123Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T10:07:39.554999Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T12:44:40.165250Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25961d7c-7140-45db-844c-b65f30441bcb · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bbc3474-3d89-46b1-81b0-83b7fa8e909d · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c392379-317a-4249-853b-1267ffc0f834 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed0f3391-e43e-45a7-8914-0450d9384e14 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning KTO: Model Alignment as Prospect Theoretic Optimization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab189571-d92c-4e31-a563-b8f435f50c8e · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Addressing function approximation error in actor-critic methods
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7bcddd2-4d4b-4d9e-9b28-5725af001341 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Reinforced Self-Training (ReST) for Language Modeling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5421bf75-b042-4d23-904e-9e889de9bd6d · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9a4d2fc-82d9-45f8-b5ea-511557d7374a · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Soft Actor-Critic Algorithms and Applications
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T21:38:16.985704+00:00.
Observation 6d86d9e5-fff9-4d18-960b-482f72409ad1 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a230ab0-e798-42e9-805b-69b896b14f7b · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Let's Verify Step by Step
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2aa0290-cb5c-4d89-a61d-df36ce13886e · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Continuous control with deep reinforcement learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aae992a3-4e35-49dd-aeec-e15126a143ae · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f098e854-5555-4e31-88ba-005aabf4d6a1 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Asynchronous methods for deep reinforce- ment learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 741670e6-7972-4a48-bedd-1e53fa95df16 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Playing Atari with Deep Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 92ac8461-ff0a-4f85-b3c7-8770b39c976e · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Tiny-grpo math tasks dataset
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ceaee7c1-da07-431c-ae74-1abfe8fef9fb · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c6d88f8-2e9f-492e-840c-505b8c086ef6 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Direct preference optimization: Your language model is secretly a reward model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b90aaba-af82-4e8a-91b9-25e702432640 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Trust region policy optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3a4c5fa6-4170-486d-839f-aa3baebb32a9 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Proximal Policy Optimization Algorithms
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 098193a6-8410-4310-a8a4-e65b507792d0 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16cf6ee6-9cb2-46a2-990c-9b69aa0e85bb · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Policy gradient meth- ods for reinforcement learning with function approximation.Advances in neural information processing systems, 12
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b21c3ea2-5585-4556-b2ca-62fa64fd7165 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Mujoco: A physics engine for model-based control
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2e1eab8-a5dc-448a-9fcc-ecfb27978475 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3a1261fc-ebbc-4e11-8e24-588501fc336d · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning LLaMA: Open and Efficient Foundation Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e6a0e3e-d178-4ce9-a48c-e126da7efb4c · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Deep reinforcement learning with double q-learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a369a7f-6f62-4214-ad9c-d5e80188561f · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Aime problem set: 1983–2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3a51b174-adb8-430c-9b7c-fc2801e4f4fe · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Multi-intersection traffic optimisation: A benchmark dataset and a strong baseline.IEEE Open Journal of Intelligent Transportation Systems, 3:126–136
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5100b9e-2bab-498c-ab30-b5f518f8c777 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Soft expert reward learning for vision-and-language navigation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ddf5a0f0-c074-401e-b154-fe352a500b3e · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Dueling network architectures for deep reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7cb80a5-9253-417f-a78f-1b3c45744716 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2df3508-e3b9-4a0e-ba58-aab5762e3347 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning fixed value
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1cf9a237-883a-4fc7-abb5-459b8cad96c9 · outbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning type": “Algebra
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c11dbf0-e383-4487-848a-3fc79f54b69c · inbound
K-Score: Kalman Filter as a Principled Alternative to Reward Normalization in Reinforcement Learning Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9146da6-0b66-425c-93b1-55f60eab6a53 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc0b0d50-775a-4cc3-95d7-dfcc27783808 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 80214a48-80f4-4075-b238-ddc558fa6ac7 · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.