Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:54:32.621093Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2605.30789.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:54:32.621093Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e5184e25-d6fb-4bbd-a786-56bd3bfbc711 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO H., Gendler, A., Baruch, E
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfdad6d3-10ef-46a2-bef7-a27be517981d · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5678be66-f935-438d-b473-6875f7e76362 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cee07469-cbc6-4be3-ba94-bfd7660981d9 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cbcfb6d9-d078-4531-8dc4-62f770210d5b · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO InternLM2 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbdbb877-69c4-4293-960d-ff94d5e1a27a · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO arXiv preprint arXiv:2505.09655 , year=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6dbc338-87b0-4128-8cd4-667a5b19763e · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Jackpot: Optimal budgeted rejection sampling for extreme actor-policy discrep- ancy
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b0bb883e-4a60-4bb8-bfc1-db939bbe8524 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Soft Adaptive Policy Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 34c1a450-5466-497c-be0a-2b225c0a4981 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Minillm: Knowl- edge distillation of large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ea111c-6918-4653-8877-3d35b65f6623 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 877c50ac-8167-4cef-bc55-0d4066c5986e · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e65d567a-c42f-4869-a487-3b2fea9c97ff · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9cf8a0d8-343e-48f9-8834-0ff07d2702ba · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d32e68f-3b47-4b32-8ab0-b5f12967a60d · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Measuring Mathematical Problem Solving With the MATH Dataset
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f8108edc-e84b-4b92-abd4-9be625991368 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Distilling the Knowledge in a Neural Network
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2f1e0915-0195-4718-ac64-e0c8493060fa · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO ORPO: Monolithic Preference Optimization without Reference Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30cc35bc-d992-49a3-977c-09b4320d912e · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Qerl: Beyond efficiency– quantization-enhanced reinforcement learning for llms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5a3bc79c-67c2-4a6c-9464-b222eb19007a · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b75fb654-4b5b-422b-90b6-17d0088be00e · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7465628b-7cca-4c32-8f66-7c8684b401f7 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Bridging Offline and Online Reinforcement Learning for LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a4016e76-460e-4422-a3b5-7515261b4ea6 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4561501a-a758-465f-99a2-04e64f88cb82 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation acb52fa1-ebfd-4aca-a3c6-894ff4fcae75 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO N., Baker, A., Neo, C., Roush, A., Kirsch, A., and Shwartz-Ziv, R
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 71d18420-4c65-4d7f-b1aa-5832cc82b84f · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Proximal Policy Optimization Algorithms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f09f70c8-c90a-4d02-8d65-37e642f77bc8 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 41debccd-ddd4-44db-af2c-9b54ce718dbe · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO HybridFlow: A Flexible and Efficient RLHF Framework
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 807a29ab-7504-4fbb-90be-3cbbad3e2165 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Unchosen experts can contribute too: Unleashing moe models’ power by self-contrast
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b14d4f-91b1-4a2d-808d-11550e860520 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b9f45e6-bdd6-4685-9e04-d696d82f0e07 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf9d075f-f5e3-4959-85d5-4501d0fd92c0 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO and Zuo, C
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d573ebba-a89f-4ef3-9e76-5f893d5e3882 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f290c786-bea9-40a4-8232-55794a183199 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Group Sequence Policy Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 71cbf156-beb6-406f-a1d3-0c7dea7708d2 · outbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Exploring multi-temperature strategies for token-and rollout-level control in rlvr
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
No inbound Pith citation observations are available.