Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:46:45.723165Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 3 of 3 outbound references and 4 inbound Pith citation observations for arXiv:2602.01460.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:46:45.723165Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T00:59:20.790604Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:31:26.842069Z
3 of 3 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 989023b2-ccb0-466c-9ad7-009b3cd9285d · outbound
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator Plugging into (68) gives E[∥ bGℓ∥2 2]≤TE h (2∥Σ−1/2 ¯ε∥4 2+2mT)R(τ) 2 i
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6944c20-4bca-479a-9fd6-f58f1b09d5e7 · outbound
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator A Closer Look at Deep Policy Gradients
Reference 666
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349366c5-ab36-45a7-8f4d-0532690b3ae4 · outbound
Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator Our setting involves a non-homogeneous polynomial, hence we state the following lemma
Reference 1978
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa37624c-5c3f-40fa-be20-6c905fd5360c · inbound
Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5a80f587-4a4f-4678-ae27-e87e20a5f66f · inbound
Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e671bd93-758a-42ce-a38b-ce5b8a065f45 · inbound
Quality-Aware Exploration Budget Allocation for Cooperative Multi-Agent Reinforcement Learning Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b8f2694b-9a50-46bc-a4da-454b9fd86a33 · inbound
On Training in Imagination Non-Uniform Noise-to-Signal Ratio in the REINFORCE Policy-Gradient Estimator
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.