Pith. sign in

Paper Citation Record · LEDGER

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

As of 18 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2506.04746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04746 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:40:08.061948Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:29:29.074324Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab798cd3-cf67-49c3-b187-b70b5b019c84 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.139524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.139524Z digest=sha256:eadfa3bbf892f58de33a404a4197455bf7d2807eb2fb8f15412e4d0342afc81a

Observation 6b1636c3-4548-463d-82af-fbc33ca6f9f3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.185507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.185507Z digest=sha256:df99b33cb8c002c4e3e2206bc8b27b2b0c2b850bf7f8d1bea0cdaa94ef22acb0

Observation 0b6f6705-fe19-4efb-9329-762b5281f9e1 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:09.135642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:40:06.253357Z digest=sha256:795880b05195959f9289be2fc17168c220a3e84d032cc533a9f82a761e23c7e7

Observation ba1f84a0-4c2f-4da3-bdc4-3521f882f266 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.320189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.320189Z digest=sha256:576432c218b62ca4491d1f1c73d127b756fd8774f660c70ef4182da21c3c03c4

Observation c8d4b795-a3c2-4ee4-843b-f256d021add3 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.380688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.380688Z digest=sha256:5338588fc23c1fe5acfd48d0c43615c346dfedb91537e9010aed6a28a6b9fbb4

Observation c89b90d7-cd9b-4837-a7a6-408fc3933a32 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.447800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.447800Z digest=sha256:84f7c9bf35677cd932c0b159e576da7937294d140b2c6361f59feac8686e0c74

Observation 64bf6571-ca46-47fb-af1f-de4e316902a8 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Training Language Models to Self-Correct via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.568303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.568303Z digest=sha256:945a1e5ba2246a2937601957772153467e617486e5b845c0f1863590a9fb0344

Observation 29ec0150-9ffb-45e1-a693-8ccdfb7c255a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.692922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.692922Z digest=sha256:8ac356fc373ebd8b4817529c6305d06daf04cdaa4a6bdb3933f6e81e0fb3ae51

Observation 853a6945-caa1-4574-bb72-90775fe66610 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Scalable agent alignment via reward modeling: a research direction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:06.800735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:06.800735Z digest=sha256:bda507582aa865f4df729bf915662487e9d1a34f4b31327b5be1d0da1851351f

Observation b5c80abc-02c5-4f3e-8a92-2b0f1e56b5d6 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:08.808434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:40:06.911148Z digest=sha256:580f11552839002fc3ebc55c1f5ff6d709215da2b62c39c40494da8ace8bf5de

Observation 0bef0996-19a8-4213-870b-8946eada23d3 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:08.480800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:40:07.035824Z digest=sha256:ad7941c9238ac32734dbd582f78d72504427b9a934f416877b8cfbac16ef780a

Observation c4301d10-da22-495e-8b85-527491d0e9f1 · outbound

This paper cites Let's Verify Step by Step.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Let's Verify Step by Step

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.148178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.148178Z digest=sha256:ab0e388d2558b7997b01455380c3e6f2b254bf72ef422fb1b928f8a8ff8e3dd9

Observation 193424a8-b370-4cb1-a1ae-fd8716f02399 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:40:08.354921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:40:07.273974Z digest=sha256:714be1eb55dc1f7d2b69f0cbe50015ee6b13ff973be29d23822bd762fba040be

Observation 68d06202-fbea-444a-bbc7-f76a46853251 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Proximal Policy Optimization Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.331186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.331186Z digest=sha256:366b6a0e8c543f4b03dc3bb9b942a1da59a87c4740a181c5a2627849fa044947

Observation 7c3ecdb3-18ab-4240-b43d-d400d0aeffeb · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.387674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.387674Z digest=sha256:50eb7d64f3b010b5ef246c31041fbb8d8b07fe036930f85488173d8da1546b78

Observation f5ad3a7b-2e7f-4bb1-9c07-ad54bd22ccaf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.456498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.456498Z digest=sha256:efd905d050e7163b4fb1291426fcab9e7b8af5f576efa3ad3afeab5aa72abec3

Observation a660bc98-c428-490f-8323-56b45bdf5715 · outbound

This paper cites an unresolved cited work.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.548585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.548585Z digest=sha256:8d5b6a49c2ff222b65fed80772ee7ff15c5cb1bddc1f31d9422cb906f01c88dc

Observation baea495c-e9ec-4df9-80e3-68e6d80d0c3f · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.604698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.604698Z digest=sha256:b08160b66d1a71ac6e6b3e9e403ecf0e97e1d2b17913663a0261ac067ffcda8c

Observation 4998d84f-6eec-4663-a3dd-345085d855e6 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Solving math word problems with process- and outcome-based feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.664835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.664835Z digest=sha256:d38bea4f7c0b8896587dc4ea2c1e6ff3eddabe9b6151bf1b6bc6f1625b285299

Observation a2c895e6-e531-4eac-987d-fceec4977b1e · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.733765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.733765Z digest=sha256:5e330b54cef28ba43a8c9501a698ba24ed06480d07acd4d46b22eb25749831d7

Observation 98430df0-679c-475b-be78-8a36753bf931 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.799079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.799079Z digest=sha256:141a4c9576b904280691ef1f8a16a937cf53f717046430b483dd6dc5fea08a29

Observation c6288b23-6d7c-4ca5-9ea3-4f13dee11fe3 · outbound

This paper cites Free Process Rewards without Process Labels.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models Free Process Rewards without Process Labels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.886754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.886754Z digest=sha256:29ed5c8a9683f9e56414b92c34ee4be32e42b87a6fbc7e617f8e31dbcadbd305

Observation 34338a91-b7fb-4e61-85dc-faf337e5ba63 · outbound

This paper cites online" 'onlinestring :=.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models online" 'onlinestring :=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:07.990278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:07.990278Z digest=sha256:ca7a479b30b85f06f9bcaa6a78cdd8a1e3c99229616121c65ade7de7427676a3

Observation ad824c78-4743-4b58-8b78-7ffe378e44a0 · outbound

This paper cites write newline.

Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:08.061948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:08.061948Z digest=sha256:d00815f671625c6b0f319f292dafeb7a3645c0ae9f5af56131725110da58a0a3

Pith citing papers

Observation fe35ccc6-759e-4a2e-b59d-15c4c08ee3a1 · inbound

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI cites this paper.

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-02T11:29:29.074324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:29:29.074324Z digest=sha256:743d8add7ffdccffd96e65f744b912b07ff81f16155471c1033d6a8d2fa13a42

Observation 6f8be023-0867-428d-8c04-a4f53cf464fd · inbound

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction cites this paper.

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-30T22:49:42.994634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T22:49:42.994634Z digest=sha256:36cdb096edd9fc069a0b5b328f3cde061302d89a2ab7f808577dcc0e36ff106c