Pith. sign in

Paper Citation Record · LEDGER

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2506.04746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04746 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:40:08.061948Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:29:29.074324Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab798cd3-cf67-49c3-b187-b70b5b019c84 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.139524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.139524Z digest=sha256:7d77b6733d32ada10b81ef9bd748d06a2e4150d61dc0fff0b6f2ba5adbf3b5ff

Observation 6b1636c3-4548-463d-82af-fbc33ca6f9f3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.185507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.185507Z digest=sha256:c1c4d31b672f7f13e8d9fa27bfb9e53de1fe7577e306b1eee9abbab1d6f958ef

Observation 0b6f6705-fe19-4efb-9329-762b5281f9e1 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:09.135642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:40:06.253357Z digest=sha256:411af1eb2a7fde5164fdcb33425814a86466dcac82cf5b41d066ad49b0e81ee4

Observation ba1f84a0-4c2f-4da3-bdc4-3521f882f266 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.320189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.320189Z digest=sha256:93e70e1181f93261c9d6c89e8c26a7ae7a51772c3483a8ecd3d75246bd045214

Observation c8d4b795-a3c2-4ee4-843b-f256d021add3 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.380688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.380688Z digest=sha256:62b2744302aac689d5fcb8e481c89803565d5efa16c1b8bd4b367fed6dcddfb1

Observation c89b90d7-cd9b-4837-a7a6-408fc3933a32 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.447800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.447800Z digest=sha256:52eaee2ce93a95c1f294fad2743dcebba4ad3f27c3a9d7b3577cd37f12dd5a06

Observation 64bf6571-ca46-47fb-af1f-de4e316902a8 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Training Language Models to Self-Correct via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.568303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.568303Z digest=sha256:20b8e3db8a79d9e2c7688d317bad3fc76d7a27a99acbc919ddfa77ea7dc5ad7e

Observation 29ec0150-9ffb-45e1-a693-8ccdfb7c255a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.692922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.692922Z digest=sha256:799088ed36a261fc7872bbfe04aec60966b79c2ed08f0043f0e4a2ee123d3d13

Observation 853a6945-caa1-4574-bb72-90775fe66610 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Scalable agent alignment via reward modeling: a research direction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.800735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.800735Z digest=sha256:03651e7beae59857727a3b1d7f78d815b2d7683ff3880281597ec2b9faeb3cc4

Observation b5c80abc-02c5-4f3e-8a92-2b0f1e56b5d6 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:08.808434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:40:06.911148Z digest=sha256:d772c83ae47190fbfd1065199eb39d7acfb7c77842207b25596f1d079c444ea3

Observation 0bef0996-19a8-4213-870b-8946eada23d3 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:08.480800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:40:07.035824Z digest=sha256:3474d9dd22a1ac0a8caf710ce58c0fddbcc3c1b8f7f6001551508e910ac90d8d

Observation c4301d10-da22-495e-8b85-527491d0e9f1 · outbound

This paper cites Let's Verify Step by Step.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Let's Verify Step by Step

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.148178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.148178Z digest=sha256:e8f8eac792b1f6b56ed66a3a16cd7e4f9b360acbc08c5c2ddab8a831f477d36c

Observation 193424a8-b370-4cb1-a1ae-fd8716f02399 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:08.354921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:40:07.273974Z digest=sha256:d4ee0e1a0a2b015fe2331debc56eb9b726342eb1caa4b008ca042015a9fda3ec

Observation 68d06202-fbea-444a-bbc7-f76a46853251 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.331186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.331186Z digest=sha256:b77e13d9908e222216212313fbab67d7a92233c544fe86c1cfb6cf763872fb60

Observation 7c3ecdb3-18ab-4240-b43d-d400d0aeffeb · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.387674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.387674Z digest=sha256:aeaad60b557cea0ce452762baaaf69e3247e5e83fe3eaa89160b1fad2a825896

Observation f5ad3a7b-2e7f-4bb1-9c07-ad54bd22ccaf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.456498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.456498Z digest=sha256:457e773fe5a28f6700019d57ec8dedd94b26a7a8a84c66311094709aea7091fa

Observation a660bc98-c428-490f-8323-56b45bdf5715 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.548585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.548585Z digest=sha256:7008f677c9c9766fb989b7deaa7a207416cc0c7f8c33ee9eea35d0f081c44057

Observation baea495c-e9ec-4df9-80e3-68e6d80d0c3f · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.604698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.604698Z digest=sha256:73afd8d4c5d24da342916093800b7f43e3a3b151138d444e0d1f30abf9ae786f

Observation 4998d84f-6eec-4663-a3dd-345085d855e6 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Solving math word problems with process- and outcome-based feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.664835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.664835Z digest=sha256:4620d4f244f951520fc1b986821cbf56739e472da9f168f8dc31510725908b7a

Observation a2c895e6-e531-4eac-987d-fceec4977b1e · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.733765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.733765Z digest=sha256:e0110086c72f729517abb890b22f683c26d04ec44d0e3efdf79eb5819b608327

Observation 98430df0-679c-475b-be78-8a36753bf931 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.799079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.799079Z digest=sha256:23805ef47ceae7ba4a4e2719ca9d51b98461711ac18721c1e65edc84622b2a55

Observation c6288b23-6d7c-4ca5-9ea3-4f13dee11fe3 · outbound

This paper cites Free Process Rewards without Process Labels.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Free Process Rewards without Process Labels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.886754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.886754Z digest=sha256:6cd7403ceb5fe302855da5a448957a3c495d750865a308e1f2845beb0b7c3f48

Observation 34338a91-b7fb-4e61-85dc-faf337e5ba63 · outbound

This paper cites online" 'onlinestring :=.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models online" 'onlinestring :=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.990278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.990278Z digest=sha256:be1a205e0524da0b24a75421f7aacb7e562e7fbb62bfce02ece6eecf57375310

Observation ad824c78-4743-4b58-8b78-7ffe378e44a0 · outbound

This paper cites write newline.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:08.061948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:08.061948Z digest=sha256:843f681000a516e3016931966ed4c5f8236cde2d403eae11e955539ad47c8282

Pith citing papers

Observation fe35ccc6-759e-4a2e-b59d-15c4c08ee3a1 · inbound

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI cites this paper.

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-02T11:29:29.074324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:29:29.074324Z digest=sha256:369b10cd8d2856ce996bcb693c9234d1802bbbe4571d3f42d9e0319b9a3ac1d0

Observation 6f8be023-0867-428d-8c04-a4f53cf464fd · inbound

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction cites this paper.

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T22:49:42.994634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T22:49:42.994634Z digest=sha256:29b73aaa279b6b4d78d6d44490ef371157133f98de54bd2a9890e52ec0fc8a7d