Pith. sign in

Paper Citation Record · LEDGER

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 7 inbound Pith citation observations for arXiv:2509.01321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01321 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:45:41.982954Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T14:49:39.188220Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.035824Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88bfa3fa-7790-4cb2-a1f6-177f484a5ec9 · outbound

This paper cites This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi).

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.198696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.966040Z digest=sha256:246a3a172252bf3072ce72e6b1f534afc3dc31a5c82374e946763466fc8de171

Observation 5a34c283-42b9-4ecd-8940-9542f172ac93 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.919862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.919862Z digest=sha256:028614d20a85aebfff507ae84bd5acb601870e744ccd1390173dc6c6eed77a69

Observation c075a77f-a5b2-4ff2-bfd0-c722d676654b · outbound

This paper cites Determinantal point processes for machine learning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Determinantal point processes for machine learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.926767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.926767Z digest=sha256:8eb072a71909721a7dc2b3b66c46ef3a6d0c8abdff3d7cfa4b8cdf450a90685c

Observation c9397d96-dbf1-4a87-929b-45ee2b2021b6 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.937712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.937712Z digest=sha256:1eab791bd6fdedeae1178a791650826d291e97bd72261e1c4ac117544bca824a

Observation 0bd7a51a-db29-4631-b72a-09e361584711 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.944746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.944746Z digest=sha256:20ba114f31f392e91341adda0b7896eb5aff427f9f2d934e9cd8bd1c3c418750

Observation d1159417-894d-428f-a0b1-3afcd9f06bff · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.948669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.948669Z digest=sha256:ea0cd5157b86a189d46183ee20aa398090a49d38ef3d346edcbf1b999a78a825

Observation a9dcc236-a80b-4850-bcbb-1a067acc5307 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.952290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.952290Z digest=sha256:59c62acca56f3d717eab4ee4b88f0c72d04faee84eccaa144b9caa20cfd497e5

Observation 43746aa5-ec9c-4df8-8e5c-4f0de72a7011 · outbound

This paper cites HARP: A challenging human-annotated math reasoning benchmark.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward HARP: A challenging human-annotated math reasoning benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.956023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.956023Z digest=sha256:6544dd0591a5829db1cccaac785f42625259ffac1ad801990cbe00fd243ae690

Observation bcf8172e-f661-4089-b7e5-131ca1103dfb · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.959563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.959563Z digest=sha256:eb6cfac55df1a845e2525e657f1e3d923c2a6b855f1871b8643651385b295b4e

Observation 5e0a894f-8988-445e-8468-2d97eeda04ba · outbound

This paper cites A Survey of Large Language Models.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward A Survey of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.962651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.962651Z digest=sha256:815cd12af4cb01567f1a07481894bf99d46639d571ec5078d680a335c39d3da8

Observation 48f3f476-4dc4-4040-9d57-7bfa4cf05fbe · outbound

This paper cites We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.188191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.969785Z digest=sha256:731acf689c08a63f6b303fe9bff3c6b2f09d6d3ede10670b128ebf6e0cdc9c82

Observation 6e6c5ecd-519d-4fa9-952d-eb28c627c9b0 · outbound

This paper cites We follow Zheng et al.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We follow Zheng et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.158494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.979652Z digest=sha256:53ba0289b5e5a75a7ba1498b396f14142fbedb25dc92478a6a5ffe4e5ec70743

Observation c3b2bdac-a2c9-44be-987f-c4d4608ca16d · outbound

This paper cites • Random: Randomly samples data from the training set.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward • Random: Randomly samples data from the training set

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.148614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.982954Z digest=sha256:580a1ced1b2db46e5323ee5095050bab8dceec5a24510a6b9aeaa3b58fd0e6eb

Observation ceba9f09-7548-451a-a022-897a36409df8 · outbound

This paper cites Similar to Yue et al.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Similar to Yue et al

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.178162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.973389Z digest=sha256:2f1efda07ec247567fb3c9aed68265f7d1df880d941266ff3e6ecab6e1fc6281

Observation c3b6a25e-7df7-48b1-9668-7093fca649c4 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Reasoning with Exploration: An Entropy Perspective

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.916058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.916058Z digest=sha256:d8e585f082dff6e092a68f8cf8f7d6ddc12fb7409f607ad971f12428c6d67efc

Observation 05558ae2-4940-4aa7-8ec2-355ffabdf59b · outbound

This paper cites User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.168425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.976582Z digest=sha256:4d6de7a1668a2809dea0fbaed96c87bed623c9bc3e4e3bf9c19acbd81414f084

Observation 5bda1edf-bfb7-4388-b30d-6e23b180f3e0 · outbound

This paper cites OpenAI o1 System Card.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward OpenAI o1 System Card

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.923179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.923179Z digest=sha256:3c4ae86a47a431890498c5acc2dafdea3dc225f991bdd25274e03afffaacdf82

Observation 2940da6b-fcee-47ff-9612-fdcb03b71617 · outbound

This paper cites LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.930273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.930273Z digest=sha256:d9c9f04b5aa5bfa55ef1555b77b1f60e1001b1333248658e97943103dc73877e

Observation 6deb6c8d-754b-4b7d-ab59-6d44bd4c19f8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.941078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.941078Z digest=sha256:6c1dc8bf5108b286e513770f092499c3bd4339f7f2e3143924b485162416768f

Observation b113e1de-a27c-4a7e-9a9b-8bb9bd634cc6 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.933948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.933948Z digest=sha256:4e546e1273b1ee3b8be9de88f2c5e855dfb66ef18fc018e98ee42aee41bb585f

Observation fa6945b6-7936-49d9-b632-7454cdcb75f3 · outbound

This paper cites Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.208226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:45:41.912268Z digest=sha256:e0ecbcc2ae2156f2f58ee4eb753759ed126fdade6a4daf89ba7115e6a6bb71ee

Pith citing papers

Observation c7f55504-31e6-49ef-94f0-d257b98496ea · inbound

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning cites this paper.

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:53:35.117871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T21:53:30.911773Z digest=sha256:5d4bab731e12c15353dfe18c018671025db1ac065ad5a7d0de0a6739d8e879e7

Observation 0017b414-69ab-4557-ba87-51182b5626e4 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.312129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:bb62082cbec0b4f849bc98bec2c8862fc438b3237a728a6a94c4fd4ddfdf9822

Observation 3a17c0a3-cb51-4d50-becd-eea6adcefb54 · inbound

IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage cites this paper.

IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:53:31.190838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T14:49:39.188220Z digest=sha256:6331dd617db601eafa70789e62da86de474333180a1f649a43face6e6fe02c01

Observation 30d0e72f-8d3b-44f4-9bf3-76c60791f499 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.366003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:13f0d88a21af8a61bcf6f7c7f170cf615b766b780832512fc5c4b773caa5b88e

Observation 0d6c5df4-001e-4297-abea-637ad4718b16 · inbound

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots cites this paper.

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.260794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:55:09.927034Z digest=sha256:ac5556920902628936f543800cb99aa8fd10befb7cebb2828e045ca0afb071ed

Observation 642a7c29-2225-48ce-a3f3-600d24e1f83a · inbound

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short cites this paper.

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.544665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T17:18:03.459289Z digest=sha256:ac7fcad8756cebed9386939a741673eee57b355739d95c507ec50c75aa8cd01a

Observation c639ed77-de36-4ef0-ae45-59cbff164918 · inbound

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models cites this paper.

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.037593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T18:29:47.613424Z digest=sha256:36c39733090f97ebe14a2c9c69c8b3d08439ba5158b982ca15b969113d4088d0