Pith. sign in

Paper Citation Record · LEDGER

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03119 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:57:30.540771Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecba873f-4762-45c8-81f9-16fbd33de8e5 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.386244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.386244Z digest=sha256:d7d43854404df56cbdfc59b661ea378e42b6a755346e7a9a483aea93d5f419ad

Observation 07b09dfb-23e2-4526-b328-55a6c87417bd · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.390868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.390868Z digest=sha256:a7d09957d315715b40f6fcfef5d51071ef2cb3a7a6d41f7851c4b4cd8eaf2fed

Observation d87970c5-f31c-4c38-a7f7-36f9bb00f687 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.394592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.394592Z digest=sha256:2b413721032c5944bcf8976b652d914e06af41a649c88510461bb578fa4f51ae

Observation 18005705-c99c-4694-8124-143e4fa47855 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.399182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.399182Z digest=sha256:e646c854de32fddea2cab103ae2387629cc412e62df1383d410b8e02209caae6

Observation 97f01257-2ca2-47b9-901a-e1134505b677 · outbound

This paper cites The Llama 3 Herd of Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.403611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.403611Z digest=sha256:d4c959ec4676e8c7f2a2359b455ed36c98232021e9250ad51e732e701472d700

Observation e80c2880-744f-4faf-8eb8-199be991df52 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.408060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.408060Z digest=sha256:ea0931659bcf703453fa7bc5cb0bdbc3600c046055eba2c0013fe9dd754c4831

Observation 74b9c92d-b17b-4834-8392-bbf267c0b8a6 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.465507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.412971Z digest=sha256:cf30373b06195c291329d1d7c7534d0d05f02fb965ffeee670c94273bfe40257

Observation 484fb415-0e14-4aac-8b65-118275eaffee · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.416789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.416789Z digest=sha256:b6138b80213bfc285afe22c5653d2547be84a0462d44a74ea4e99069ddd56baf

Observation 31830e38-500f-4397-aac4-55b103e5e2ae · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.454145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.421426Z digest=sha256:03fdd4f7d1310e7c8f2c065ed916234b4a491a9be486500bac2cf4c2cd368f27

Observation 99c47dc0-d0af-4189-87b4-953c5e17baf9 · outbound

This paper cites OpenAI o1 System Card.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.425198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.425198Z digest=sha256:28338a9516311c239f576667461fc364f5dd279590f43d09456e8116bb818654

Observation 4d683406-7572-4bfa-ad39-ffa84c38e2fb · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.429284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.429284Z digest=sha256:75e2ebdbbd376a4ed81246fcac3ed6ee684893659b9481d796db3bf62fcf90fa

Observation 8a4d22af-d5c6-4369-b5b9-1450eff70389 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.433618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.433618Z digest=sha256:7db7b571aa1a8aff12e71e56ca2efe2b266658dbbf070fe441cd77793eb8ded9

Observation 856f1773-7b5e-4e2c-af7e-5ab9a838f22f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.437630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.437630Z digest=sha256:df445aa3289c6b3c5aece5fddd284ea7230c5d38e26df64678188f3e6d5aee82

Observation acccc0d8-e472-4e29-b736-14e9ece3952e · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.441747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.441747Z digest=sha256:dc00110bd335553f3a3d8807a4dbd0d92a82dd6a25bb60831d960291928bc0e7

Observation 6a4a8b41-64b2-4fe0-8847-50fbe8d8f044 · outbound

This paper cites Let's Verify Step by Step.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.445813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.445813Z digest=sha256:5be197e705f1dd6baaa6892bacb3e4021c02b7f4909c05bef693a5461332f0e3

Observation 639fe561-f601-4de2-9b25-cb60c4cb8fcf · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.443605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.450080Z digest=sha256:17087bcd73c9b0dc3bc8e2aacc3df75e4d6fb8f7711027ad34a116b2098c347a

Observation d9b300f8-7a82-4088-a60b-468e9679e659 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.432998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.454210Z digest=sha256:d139319844f8e9a9232baf33774d8e81d859dc0f56b8c8b3206ce81a5f598735

Observation 031d6294-c5f6-4f4e-848e-3be310f62c2b · outbound

This paper cites Training language models to follow instructions with human feedback.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.458018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.458018Z digest=sha256:db6a6ba044bf0e3ce57cb6836c79dc28985cb6b1d98118f037ff419eb0dd0bb2

Observation a1838133-bd51-4c1f-bf27-58ee53ce2e5c · outbound

This paper cites Language Model Self-improvement by Reinforcement Learning Contemplation.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.462207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.462207Z digest=sha256:773e6f21732140efde331f666c3a25f53b3f86377918d0790eb431ac0597ae53

Observation a5c27ee9-18a9-44e4-a8f8-2f1ba7aba2c1 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Maximizing Confidence Alone Improves Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.465960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.465960Z digest=sha256:1c3578d35739762f1208c90fe181baf579f1af3ff778a07e9523def1fd9447f5

Observation e4444bbf-596d-4703-8f07-b06a47a936e5 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.469657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.469657Z digest=sha256:69238212f9841599ceb60649310ca673bd8c3174864fd983475afe561d6bf2da

Observation 2407f428-d323-4c52-b40f-01f50de7819f · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Manning, Stefano Ermon, and Chelsea Finn

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.473324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.473324Z digest=sha256:63c8bc01d024c39c240fdaa59aa68b9c59ba9b07fa92b9640f2df4c92dc9df86

Observation ba4562bc-c250-46e2-b7e3-d20bcea45c59 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T00:57:30.854553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.476769Z digest=sha256:1dbe7077491de237674b9139bd8bcffbd8e1a8cd74353f99e20499aa53689f48

Observation 4ebe5c2a-8a41-4d26-96b6-4460b9107f1b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.480119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.480119Z digest=sha256:93cf6472d738eb5890d7adc8efcc6c4c1ef845bbc47476bec0fd9e7c5896b599

Observation 6e233841-6b2d-4913-abe1-03758e78e6cb · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.484139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.484139Z digest=sha256:fc3112032b353fd0b264005254a9992e668324d79f212332695f47d827a492b3

Observation 52e9b7ae-8a61-4a64-b77f-673325afb5da · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.487540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.487540Z digest=sha256:7dbeb8cc432e80c99a5b8c30a3a40357be2b006fedd5ca9cbab86af192a2db95

Observation bd697cd9-bfbe-4342-a776-d9aa8ca12e60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.491054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.491054Z digest=sha256:7397386567d917977aad0d58bfa3eb25e2c5927caa01d4b3f89934be348b6f1b

Observation bffc2168-0b82-4e2f-8f03-a67983067e19 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.414598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.494817Z digest=sha256:a2347ef1ac7ce20f603540f7b40102b211d4483cb11fe9bbdf6450b9b58d0c94

Observation c19251bb-e489-49b7-8ef3-3a65dfaf0ad7 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.498471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.498471Z digest=sha256:6c50266735dcbe8f4ff2e68974387c2aa365df5e965b5145fdcc0b28a10985d2

Observation 0b698316-a173-408c-be4d-e7d77d98a12c · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.502123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.502123Z digest=sha256:fed62b33725a081704b0a2b5d660ac628c9c2a84a6f5ca951b252ab947c720f8

Observation a6584d1b-77ba-405b-814e-f56b6451ac95 · outbound

This paper cites Qwen3 Technical Report.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.506098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.506098Z digest=sha256:c0f0b26b3e364837607a564d76be30403980fb5fa295e7c2130394897b9b1fef

Observation f2b0c7ad-135f-4a0e-90bd-f0ec6b1178b7 · outbound

This paper cites Qwen2.5 Technical Report.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.510294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.510294Z digest=sha256:0dd138f6b4ba5e8dbcfc6e4f2b46d0963a3fcc2100e30985600c910776f844bd

Observation e31dbe6a-4007-4b9e-85c6-03f5a818230c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.514403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.514403Z digest=sha256:068a1a3625dfd5331e66f6fedcade35fbe5ecb862d95cea117b0907de0bd5989

Observation 56d7b9d6-bb2d-43c0-819b-d3f6c8771485 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.518105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.518105Z digest=sha256:a5f25ef1c852194c3021a4cbfec440e12e733db155ab8285a207a5a25a8da57b

Observation 71176864-784a-4c8d-a116-44d3b33deb16 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.521694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.521694Z digest=sha256:a3dbd5c1cd26a77f90984b11a7f44dc73ba0041393f6f1a6d4109f465cd6c7f4

Observation 5bdbee35-1a8f-45dd-833d-2f3053cbf2a8 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 36

Resolution
verified exact
doi, observed 2026-08-08T00:57:30.666565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.525459Z digest=sha256:4144136f24f4c37c806a9d1bb749c034cb1604dbcb5e702b2c9d5a9206ace6ae

Observation 91f079ec-09f0-4ead-865f-40038c090f19 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.529006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.529006Z digest=sha256:425e8a570f61006bf35ffce415a30bc677a7c5dd07cf3a754ba84b40cefc6af8

Observation a3e3bafa-9008-4732-aef9-bfd00e0bbdf1 · outbound

This paper cites Learning to Reason without External Rewards.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Learning to Reason without External Rewards

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.532594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.532594Z digest=sha256:adca65d8c85b057524a836b87426f40e40ae30f469cebc4cab5316d7b06f529b

Observation 9b1374cf-3fea-4de6-aa10-19a0e09406a3 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Instruction-Following Evaluation for Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.537043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.537043Z digest=sha256:9c89da8990775d585d815d657b69b866a78f0a95db0975dc46c500031f8250ee

Observation df8cc502-09ba-410a-9a0b-c6e8240da32c · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.540771Z digest=sha256:0bd2d7d188cfbb05d7ad48ca3eaac2fc362f81fc6b60bc8d979892f5ee7aa3ba

Pith citing papers

No inbound Pith citation observations are available.