Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:40:08.061948Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2506.04746.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:40:08.061948Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T11:29:29.074324Z
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ab798cd3-cf67-49c3-b187-b70b5b019c84 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Training Verifiers to Solve Math Word Problems
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1636c3-4548-463d-82af-fbc33ca6f9f3 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6f6705-fe19-4efb-9329-762b5281f9e1 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba1f84a0-4c2f-4da3-bdc4-3521f882f266 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8d4b795-a3c2-4ee4-843b-f256d021add3 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c89b90d7-cd9b-4837-a7a6-408fc3933a32 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bf6571-ca46-47fb-af1f-de4e316902a8 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Training Language Models to Self-Correct via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ec0150-9ffb-45e1-a693-8ccdfb7c255a · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853a6945-caa1-4574-bb72-90775fe66610 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c80abc-02c5-4f3e-8a92-2b0f1e56b5d6 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0bef0996-19a8-4213-870b-8946eada23d3 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4301d10-da22-495e-8b85-527491d0e9f1 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Let's Verify Step by Step
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193424a8-b370-4cb1-a1ae-fd8716f02399 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68d06202-fbea-444a-bbc7-f76a46853251 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Proximal Policy Optimization Algorithms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c3ecdb3-18ab-4240-b43d-d400d0aeffeb · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ad3a7b-2e7f-4bb1-9c07-ad54bd22ccaf · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a660bc98-c428-490f-8323-56b45bdf5715 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baea495c-e9ec-4df9-80e3-68e6d80d0c3f · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4998d84f-6eec-4663-a3dd-345085d855e6 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Solving math word problems with process- and outcome-based feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c895e6-e531-4eac-987d-fceec4977b1e · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98430df0-679c-475b-be78-8a36753bf931 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6288b23-6d7c-4ca5-9ea3-4f13dee11fe3 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Free Process Rewards without Process Labels
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34338a91-b7fb-4e61-85dc-faf337e5ba63 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models online" 'onlinestring :=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad824c78-4743-4b58-8b78-7ffe378e44a0 · outbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models write newline
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe35ccc6-759e-4a2e-b59d-15c4c08ee3a1 · inbound
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8be023-0867-428d-8c04-a4f53cf464fd · inbound
LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.