Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2505.02391.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:01:49.664298Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T06:06:41.070540Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation bccae710-a182-47cd-93e1-f582b97def4c · inbound
SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41bb5f15-b3ab-47ba-9182-7d978c723933 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0549afe-c503-48f6-b347-d34c7c8c723c · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27f7bb9-1692-4506-9202-d39226132c92 · inbound
A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5063b41d-8853-4fbb-88d2-cb6914f75e65 · inbound
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e1a0d99-f6cc-429c-96a5-c48079f6e7c8 · inbound
Your Model Diversity, Not Method, Determines Reasoning Strategy Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f7167e3-4397-4ab1-b6f7-9864df108db4 · inbound
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b6111562-bf36-46ae-bf57-8811dabd805c · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 514e9eb2-f1de-49df-ad60-8c9559b9b20c · inbound
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.