Pith. sign in

Paper Citation Record · LEDGER

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 7 inbound Pith citation observations for arXiv:2509.01321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01321 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:45:41.982954Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T14:49:39.188220Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.035824Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88bfa3fa-7790-4cb2-a1f6-177f484a5ec9 · outbound

This paper cites This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi).

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.198696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.966040Z digest=sha256:afcdd96a46df1643161b8217639b90fa1b9fe745c01f740122cb7f0776a59f4e

Observation 5a34c283-42b9-4ecd-8940-9542f172ac93 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.919862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.919862Z digest=sha256:e7eeaf70d342341615daf7e4fa60253a6c477fe767aefeaeefb7ed421504dfe1

Observation c075a77f-a5b2-4ff2-bfd0-c722d676654b · outbound

This paper cites Determinantal point processes for machine learning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Determinantal point processes for machine learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.926767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.926767Z digest=sha256:bc728154104a990847c15e05aebca0bda66da41551c6844d84c1bebd4bd4f345

Observation c9397d96-dbf1-4a87-929b-45ee2b2021b6 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.937712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.937712Z digest=sha256:58079561e2f4086cca3ef17e95d558b6450f5729cc3fb83f42d6a07b108fd065

Observation 0bd7a51a-db29-4631-b72a-09e361584711 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.944746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.944746Z digest=sha256:56fa390d159b83766bf4c55550d37c13db181d48e35a54c22009ce88c5b6c129

Observation d1159417-894d-428f-a0b1-3afcd9f06bff · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.948669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.948669Z digest=sha256:08b31931159aeff3596dd511ce1cd3d6b75aef4ed613fcc4a2fc1a00ef6011b3

Observation a9dcc236-a80b-4850-bcbb-1a067acc5307 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.952290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.952290Z digest=sha256:88dc3e8d0e51cc2b20d777e78f3c834db8f4112c07101bbf4465f688d7b4ec08

Observation 43746aa5-ec9c-4df8-8e5c-4f0de72a7011 · outbound

This paper cites HARP: A challenging human-annotated math reasoning benchmark.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward HARP: A challenging human-annotated math reasoning benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.956023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.956023Z digest=sha256:9af3634ca3c1503f49a2e24e6e4cda896cf8191854531bcbe1c24653d934f557

Observation bcf8172e-f661-4089-b7e5-131ca1103dfb · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.959563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.959563Z digest=sha256:6dbd7b6fea722875587bea4dc5baae5d5122d6e9e01639a029e4d1ae3cb07d5b

Observation 5e0a894f-8988-445e-8468-2d97eeda04ba · outbound

This paper cites A Survey of Large Language Models.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward A Survey of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.962651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.962651Z digest=sha256:03147092dacda06c179cd7928516c9ff2c1ac6fb8ded60b670420ac642d66dcb

Observation 48f3f476-4dc4-4040-9d57-7bfa4cf05fbe · outbound

This paper cites We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.188191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.969785Z digest=sha256:f031dfc827a96e209c84f0b2ceeb102445eb288202840d5726d0e0594863d40a

Observation 6e6c5ecd-519d-4fa9-952d-eb28c627c9b0 · outbound

This paper cites We follow Zheng et al.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We follow Zheng et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.158494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.979652Z digest=sha256:1e40fdcfd962fb6340f9890f4f63e276a083bb79338907f31778d81e68f536c1

Observation c3b2bdac-a2c9-44be-987f-c4d4608ca16d · outbound

This paper cites • Random: Randomly samples data from the training set.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward • Random: Randomly samples data from the training set

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.148614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.982954Z digest=sha256:f717da4d3278a6432f11648fc4a0831a06f25c616b9c951472e51d9a19f56c67

Observation ceba9f09-7548-451a-a022-897a36409df8 · outbound

This paper cites Similar to Yue et al.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Similar to Yue et al

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.178162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.973389Z digest=sha256:e3a69339f873ae64f1ef731d6476ddc3e86f1d9e4007c548f78805faeb99932c

Observation c3b6a25e-7df7-48b1-9668-7093fca649c4 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Reasoning with Exploration: An Entropy Perspective

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.916058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.916058Z digest=sha256:49bd29cc1222cdac2c01e7bb29851e3d5859a7f02ff0ddf2c780178ea2f1b711

Observation 05558ae2-4940-4aa7-8ec2-355ffabdf59b · outbound

This paper cites User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.168425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.976582Z digest=sha256:eb8dea9bdd92208041260c57384c501839181caed28c3a11ac180f7bbf1af811

Observation 5bda1edf-bfb7-4388-b30d-6e23b180f3e0 · outbound

This paper cites OpenAI o1 System Card.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward OpenAI o1 System Card

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.923179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.923179Z digest=sha256:3df35064b55af077ba9f5df9e54b275bf189789f8f7c04e6d276b4a01599bff3

Observation 2940da6b-fcee-47ff-9612-fdcb03b71617 · outbound

This paper cites LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.930273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.930273Z digest=sha256:39f062ce8b4fd85043a6101e402104e95d8402a6ffeae649452d67b7a5cfe336

Observation 6deb6c8d-754b-4b7d-ab59-6d44bd4c19f8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.941078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.941078Z digest=sha256:65b0d080c128b6cfc2258582dd6c58b96fda0e7d040c979c2070bad04a0afa16

Observation b113e1de-a27c-4a7e-9a9b-8bb9bd634cc6 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.933948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.933948Z digest=sha256:dc3bbf77a8ad36e1658ddb2bf5cd91e697a92f53fcc4777d23dd27f447335d23

Observation fa6945b6-7936-49d9-b632-7454cdcb75f3 · outbound

This paper cites Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.208226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:45:41.912268Z digest=sha256:7e21838c713153ca7f3fcd7e911c85475d7608319916c7683ded09d4fee2eeda

Pith citing papers

Observation c7f55504-31e6-49ef-94f0-d257b98496ea · inbound

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning cites this paper.

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:53:35.117871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T21:53:30.911773Z digest=sha256:8c45d294a39df1239561726e6479beef868834e650d7b01f261919bf8886abbb

Observation 0017b414-69ab-4557-ba87-51182b5626e4 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.312129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:2d7f9e8b7b1ce3c8549678f04a99b21a6d366072f2d8a0021fe0ec267a4e5561

Observation 3a17c0a3-cb51-4d50-becd-eea6adcefb54 · inbound

IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage cites this paper.

IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:53:31.190838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T14:49:39.188220Z digest=sha256:b60ba67343b14f697ee28b1b21c23a385cd206cb93bd594577d32c042b3819a9

Observation 30d0e72f-8d3b-44f4-9bf3-76c60791f499 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.366003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:ce7a9a48c4c0fea45f7612078e38e174c21871922326ce444ef5d58b6f1e46a1

Observation 0d6c5df4-001e-4297-abea-637ad4718b16 · inbound

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots cites this paper.

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.260794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:55:09.927034Z digest=sha256:c40a10a195dc9f81d99c73b1d24408da8c1cba0a1554c531e1f2a84ff9962735

Observation 642a7c29-2225-48ce-a3f3-600d24e1f83a · inbound

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short cites this paper.

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.544665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:18:03.459289Z digest=sha256:0343b3943798fff261158072908b63a8c1716142b727e776fd229628a097cd08

Observation c639ed77-de36-4ef0-ae45-59cbff164918 · inbound

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models cites this paper.

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.037593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:29:47.613424Z digest=sha256:78064be8a64adc45ca7ac42eca20e5313394f323d60cfbfa695033730c5c125d