Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:45:41.982954Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 7 inbound Pith citation observations for arXiv:2509.01321.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:45:41.982954Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T14:49:39.188220Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:09:29.035824Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 88bfa3fa-7790-4cb2-a1f6-177f484a5ec9 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a34c283-42b9-4ecd-8940-9542f172ac93 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c075a77f-a5b2-4ff2-bfd0-c722d676654b · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Determinantal point processes for machine learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9397d96-dbf1-4a87-929b-45ee2b2021b6 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd7a51a-db29-4631-b72a-09e361584711 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1159417-894d-428f-a0b1-3afcd9f06bff · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9dcc236-a80b-4850-bcbb-1a067acc5307 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43746aa5-ec9c-4df8-8e5c-4f0de72a7011 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward HARP: A challenging human-annotated math reasoning benchmark
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf8172e-f661-4089-b7e5-131ca1103dfb · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0a894f-8988-445e-8468-2d97eeda04ba · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward A Survey of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f3f476-4dc4-4040-9d57-7bfa4cf05fbe · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e6c5ecd-519d-4fa9-952d-eb28c627c9b0 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We follow Zheng et al
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3b2bdac-a2c9-44be-987f-c4d4608ca16d · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward • Random: Randomly samples data from the training set
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ceba9f09-7548-451a-a022-897a36409df8 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Similar to Yue et al
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c3b6a25e-7df7-48b1-9668-7093fca649c4 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Reasoning with Exploration: An Entropy Perspective
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05558ae2-4940-4aa7-8ec2-355ffabdf59b · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5bda1edf-bfb7-4388-b30d-6e23b180f3e0 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward OpenAI o1 System Card
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2940da6b-fcee-47ff-9612-fdcb03b71617 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6deb6c8d-754b-4b7d-ab59-6d44bd4c19f8 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b113e1de-a27c-4a7e-9a9b-8bb9bd634cc6 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Understanding R1-Zero-Like Training: A Critical Perspective
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6945b6-7936-49d9-b632-7454cdcb75f3 · outbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7f55504-31e6-49ef-94f0-d257b98496ea · inbound
Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0017b414-69ab-4557-ba87-51182b5626e4 · inbound
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a17c0a3-cb51-4d50-becd-eea6adcefb54 · inbound
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30d0e72f-8d3b-44f4-9bf3-76c60791f499 · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d6c5df4-001e-4297-abea-637ad4718b16 · inbound
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 642a7c29-2225-48ce-a3f3-600d24e1f83a · inbound
Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c639ed77-de36-4ef0-ae45-59cbff164918 · inbound
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.