Pith. sign in

Paper Citation Record · LEDGER

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03119 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:57:30.540771Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecba873f-4762-45c8-81f9-16fbd33de8e5 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.386244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.386244Z digest=sha256:ba138fc72ed9e6498abae4e5a357c4ee7d2bd5434486bf22f616fb10735549ff

Observation 07b09dfb-23e2-4526-b328-55a6c87417bd · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.390868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.390868Z digest=sha256:dcdf8a5fd9b316171c581d0ae22df8cf9895b7d5463b87422061acf06124a949

Observation d87970c5-f31c-4c38-a7f7-36f9bb00f687 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.394592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.394592Z digest=sha256:5ce590ec2690c8517d78c8a8212b096f384288e0f69eba329e7f64101abc24a2

Observation 18005705-c99c-4694-8124-143e4fa47855 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.399182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.399182Z digest=sha256:220aff7e35cf71eab000bf476b7114d46b22e7cba4410f7f0d4999c9dc70526f

Observation 97f01257-2ca2-47b9-901a-e1134505b677 · outbound

This paper cites The Llama 3 Herd of Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.403611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.403611Z digest=sha256:1a28f829d7c89bba3eb6cc0edb6f52a96e6d001bc1e23e6ed3ac3b53479bb50b

Observation e80c2880-744f-4faf-8eb8-199be991df52 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.408060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.408060Z digest=sha256:fa2f31e2053697a09cd20280f90b5cef525cc1409f420b71b34502b2b0524231

Observation 74b9c92d-b17b-4834-8392-bbf267c0b8a6 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.465507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.412971Z digest=sha256:97f71cc2ec0f1c9437be876a90897b496754e3fcbc508f591637abd22afa60e5

Observation 484fb415-0e14-4aac-8b65-118275eaffee · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.416789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.416789Z digest=sha256:bd99d98e6236063307edcfe71c9d0cc7cdf9dd2261220027d38f76a8fe1ecfa4

Observation 31830e38-500f-4397-aac4-55b103e5e2ae · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.454145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.421426Z digest=sha256:9419da3dc34ffcdb514cb6d52955f373453e851ff25dd293a5a8162697b16751

Observation 99c47dc0-d0af-4189-87b4-953c5e17baf9 · outbound

This paper cites OpenAI o1 System Card.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.425198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.425198Z digest=sha256:49573d623db207c90f175f6d1417e8947a77f30e07ab6179b7cb2a9e347b57a2

Observation 4d683406-7572-4bfa-ad39-ffa84c38e2fb · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.429284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.429284Z digest=sha256:50dfecd60a8fe1b1f48a96eb633e1eaa7f74df79d6d2a943dd4780e65b75b515

Observation 8a4d22af-d5c6-4369-b5b9-1450eff70389 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.433618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.433618Z digest=sha256:53f5906c8498f2d7383cfdc4f37226044ad69cbd159ca0d7a923d3e90ab8c21d

Observation 856f1773-7b5e-4e2c-af7e-5ab9a838f22f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.437630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.437630Z digest=sha256:c4541e78149c15934395e980471531e56db861cea012403965246f1588287155

Observation acccc0d8-e472-4e29-b736-14e9ece3952e · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.441747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.441747Z digest=sha256:908afe27428e41fef4f988ae1697ced57504636bb17eb7e6a0f7a7fef67c64db

Observation 6a4a8b41-64b2-4fe0-8847-50fbe8d8f044 · outbound

This paper cites Let's Verify Step by Step.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.445813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.445813Z digest=sha256:da1fc23118a127d9e1f2d4dc065084cb2ac1f27f20b36391f2bd72911b8cd18b

Observation 639fe561-f601-4de2-9b25-cb60c4cb8fcf · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.443605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.450080Z digest=sha256:acf6c485fbfa6d4cab40352eae2ce91964b2366f1615e9754e1beb240bb9361d

Observation d9b300f8-7a82-4088-a60b-468e9679e659 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.432998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.454210Z digest=sha256:6a800909caa6613e685dc144649fe6e3f3d42c0fbb16781f700e7fd07d97d22e

Observation 031d6294-c5f6-4f4e-848e-3be310f62c2b · outbound

This paper cites Training language models to follow instructions with human feedback.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.458018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.458018Z digest=sha256:1676e0c15be9ecc8818bb84b96ae3ba0beafeaa19886ba9472911c0e8ad1b6d1

Observation a1838133-bd51-4c1f-bf27-58ee53ce2e5c · outbound

This paper cites Language Model Self-improvement by Reinforcement Learning Contemplation.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.462207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.462207Z digest=sha256:7f63d596d637d2c1837d99d8c9e96cc34d303548e142fcce5c35c94778e35d70

Observation a5c27ee9-18a9-44e4-a8f8-2f1ba7aba2c1 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Maximizing Confidence Alone Improves Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.465960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.465960Z digest=sha256:6ec89081940609620b42e6eb2261ad067598257a7f76ee8b85a2d43f61fe7679

Observation e4444bbf-596d-4703-8f07-b06a47a936e5 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.469657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.469657Z digest=sha256:d75d5a7453acfa05fe0d1bba17c91e7961bea4f19088ca37b18db0056e01a9a8

Observation 2407f428-d323-4c52-b40f-01f50de7819f · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Manning, Stefano Ermon, and Chelsea Finn

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.473324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.473324Z digest=sha256:c510990d7bfe50a9f26128e6709032be71a0423ea780a9a6e349d65e8b3f1c27

Observation ba4562bc-c250-46e2-b7e3-d20bcea45c59 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T00:57:30.854553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.476769Z digest=sha256:47d21dff06e20b66f65e032abbbe69b9a5f973462c36466df2272586d025ebf9

Observation 4ebe5c2a-8a41-4d26-96b6-4460b9107f1b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.480119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.480119Z digest=sha256:f43ca7cc45a8f9d80af50a95dbd5dfb2d204de5ac84cce0fd26ba93251f555fe

Observation 6e233841-6b2d-4913-abe1-03758e78e6cb · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.484139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.484139Z digest=sha256:a5625c1ff7581ca3f67662f720ef49b7a03f5cc5d41e79e983df1f8b4e0a2718

Observation 52e9b7ae-8a61-4a64-b77f-673325afb5da · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.487540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.487540Z digest=sha256:1fabff21bfd0705dcf57ec7aee8ff423f16e8d3bc6b63dbdfb2ac338dd96b980

Observation bd697cd9-bfbe-4342-a776-d9aa8ca12e60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.491054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.491054Z digest=sha256:39f9991b995b3f72de875b1fc55f313c2a59392842ad61fd3ff1f4cb392debf0

Observation bffc2168-0b82-4e2f-8f03-a67983067e19 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.414598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.494817Z digest=sha256:8d3640a2eefaf463418a1200c1f809498c0b32f47dfae506bb889e814f301c04

Observation c19251bb-e489-49b7-8ef3-3a65dfaf0ad7 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.498471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.498471Z digest=sha256:a53fda6614cac2b8a568ef551650f2ff63409781a2de00f974bcbf2560386c4a

Observation 0b698316-a173-408c-be4d-e7d77d98a12c · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.502123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.502123Z digest=sha256:590d202761ac36bdeacc5a996750cbaf1239d43d9036aa9e0ca2dac7c40ba69e

Observation a6584d1b-77ba-405b-814e-f56b6451ac95 · outbound

This paper cites Qwen3 Technical Report.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.506098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.506098Z digest=sha256:09f2205903cf1de5e3eedc28297e194352cb9ae0734697ab43699221b33c20cc

Observation f2b0c7ad-135f-4a0e-90bd-f0ec6b1178b7 · outbound

This paper cites Qwen2.5 Technical Report.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.510294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.510294Z digest=sha256:7b166df427dd97c6854af4b13d6a5cbaba6e24cf3c285444d3226e5fe8c48eb9

Observation e31dbe6a-4007-4b9e-85c6-03f5a818230c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.514403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.514403Z digest=sha256:c6aa4836c9a35ea91a6541d77bc610cceaecd863efb2fe92d82831d5185a220c

Observation 56d7b9d6-bb2d-43c0-819b-d3f6c8771485 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.518105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.518105Z digest=sha256:994dcedd4a445ce1cdd3a7ec85282a8bcdd51eb2453e4097c98753bc00a662d7

Observation 71176864-784a-4c8d-a116-44d3b33deb16 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.521694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.521694Z digest=sha256:01a743e323f3d61131fd28ad2f610dffce25d84b695240d0b2822d2e592b925f

Observation 5bdbee35-1a8f-45dd-833d-2f3053cbf2a8 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 36

Resolution
verified exact
doi, observed 2026-08-08T00:57:30.666565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.525459Z digest=sha256:9728ee9eb5e05982df78210807f6d5e857d04a54011f3202fba0ccd49a7ee1f3

Observation 91f079ec-09f0-4ead-865f-40038c090f19 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.529006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.529006Z digest=sha256:0c3562dc5a71f906eeeaee1eac96bd797ee0139a3fd354996ecb5fb7a62665ba

Observation a3e3bafa-9008-4732-aef9-bfd00e0bbdf1 · outbound

This paper cites Learning to Reason without External Rewards.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Learning to Reason without External Rewards

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.532594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.532594Z digest=sha256:76288d748b85bbbae399416edabf3de906c836d610a3b75aa24407632cdf2fd3

Observation 9b1374cf-3fea-4de6-aa10-19a0e09406a3 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Instruction-Following Evaluation for Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.537043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.537043Z digest=sha256:8aa9a1a919e457be7c4e2ade8f76f346e752b079a6b21dc20074d07753f75fef

Observation df8cc502-09ba-410a-9a0b-c6e8240da32c · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.540771Z digest=sha256:34ea8c2f9df7c1e92d34c1c3e64c1eb5ae3ad318466aeadb41742541d8111ef5

Pith citing papers

No inbound Pith citation observations are available.