Pith. sign in

Paper Citation Record · LEDGER

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2505.02391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02391 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:01:49.664298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:06:41.070540Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bccae710-a182-47cd-93e1-f582b97def4c · inbound

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation cites this paper.

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:49.664298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:49.664298Z digest=sha256:aafd96a50f14291e45ded5a698aa12f6cf12ae98bd988492b10ae1ded7c74779

Observation 41bb5f15-b3ab-47ba-9182-7d978c723933 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:23.430181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:23.430181Z digest=sha256:0c9e2660ad4bd978adb7ca68d0a115c7109e9345434e847f788e9ae8cee91aab

Observation f0549afe-c503-48f6-b347-d34c7c8c723c · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.900557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.900557Z digest=sha256:9a2e4c6c32159896f5186cd05a5f302bb7943339c3644afd0a3328f18dc3eed7

Observation c27f7bb9-1692-4506-9202-d39226132c92 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:38.039514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:38.039514Z digest=sha256:458deaf4260d4f58393ba125d7a68f023560155863c0b4a2cba4e31673e68981

Observation 5063b41d-8853-4fbb-88d2-cb6914f75e65 · inbound

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction cites this paper.

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:26:41.631562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:26:35.767179Z digest=sha256:d9c3ccdf112b5a6c0db4c4f03bd943b8156b5ad4028d304abee01c7367932e67

Observation 3e1a0d99-f6cc-429c-96a5-c48079f6e7c8 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:02.867497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:8ed8c61848b2c0265a149525640f8fc7376612769e844451c390dcd5c6233d3e

Observation 9f7167e3-4397-4ab1-b6f7-9864df108db4 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.487341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:2420032e7c7522034c3ca4eb95a5ea7a8709bf440ba5ef902099bc1681751811

Observation b6111562-bf36-46ae-bf57-8811dabd805c · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.072433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:eb8c9b55f98fdad58c58eb5e178573645f48b773c446b141f8e974bbf8fdcc71

Observation 514e9eb2-f1de-49df-ad60-8c9559b9b20c · inbound

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance cites this paper.

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:11.865683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:22:11.865683Z digest=sha256:4d1bdeca1dc5f939b9b8eeaaeb4b82ac2a192361f3c8f82d6cbb6f3e5843b0f6