Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:57:30.540771Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03119.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:57:30.540771Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ecba873f-4762-45c8-81f9-16fbd33de8e5 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b09dfb-23e2-4526-b328-55a6c87417bd · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87970c5-f31c-4c38-a7f7-36f9bb00f687 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18005705-c99c-4694-8124-143e4fa47855 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f01257-2ca2-47b9-901a-e1134505b677 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80c2880-744f-4faf-8eb8-199be991df52 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b9c92d-b17b-4834-8392-bbf267c0b8a6 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 484fb415-0e14-4aac-8b65-118275eaffee · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31830e38-500f-4397-aac4-55b103e5e2ae · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99c47dc0-d0af-4189-87b4-953c5e17baf9 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR OpenAI o1 System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d683406-7572-4bfa-ad39-ffa84c38e2fb · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4d22af-d5c6-4369-b5b9-1450eff70389 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856f1773-7b5e-4e2c-af7e-5ab9a838f22f · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acccc0d8-e472-4e29-b736-14e9ece3952e · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4a8b41-64b2-4fe0-8847-50fbe8d8f044 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Let's Verify Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639fe561-f601-4de2-9b25-cb60c4cb8fcf · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9b300f8-7a82-4088-a60b-468e9679e659 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 031d6294-c5f6-4f4e-848e-3be310f62c2b · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1838133-bd51-4c1f-bf27-58ee53ce2e5c · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Language Model Self-improvement by Reinforcement Learning Contemplation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c27ee9-18a9-44e4-a8f8-2f1ba7aba2c1 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Maximizing Confidence Alone Improves Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4444bbf-596d-4703-8f07-b06a47a936e5 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2407f428-d323-4c52-b40f-01f50de7819f · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Manning, Stefano Ermon, and Chelsea Finn
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba4562bc-c250-46e2-b7e3-d20bcea45c59 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ebe5c2a-8a41-4d26-96b6-4460b9107f1b · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e233841-6b2d-4913-abe1-03758e78e6cb · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e9b7ae-8a61-4a64-b77f-673325afb5da · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd697cd9-bfbe-4342-a776-d9aa8ca12e60 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bffc2168-0b82-4e2f-8f03-a67983067e19 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c19251bb-e489-49b7-8ef3-3a65dfaf0ad7 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b698316-a173-408c-be4d-e7d77d98a12c · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6584d1b-77ba-405b-814e-f56b6451ac95 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen3 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b0c7ad-135f-4a0e-90bd-f0ec6b1178b7 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen2.5 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31dbe6a-4007-4b9e-85c6-03f5a818230c · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d7b9d6-bb2d-43c0-819b-d3f6c8771485 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71176864-784a-4c8d-a116-44d3b33deb16 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bdbee35-1a8f-45dd-833d-2f3053cbf2a8 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91f079ec-09f0-4ead-865f-40038c090f19 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e3bafa-9008-4732-aef9-bfd00e0bbdf1 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Learning to Reason without External Rewards
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1374cf-3fea-4de6-aa10-19a0e09406a3 · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Instruction-Following Evaluation for Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df8cc502-09ba-410a-9a0b-c6e8240da32c · outbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.