Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 2 inbound Pith citation observations for arXiv:2606.12370.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:48.709434Z
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7a1fdafe-eeae-436e-867c-e1b7bb751525 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Draft-OPD: On-Policy Distillation for Speculative Draft Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d0b5a34-4160-4c2e-ad9a-8c1252c72cc3 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The TV gradient is proportional to qj, so it automatically ignores low-probability tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 358997f3-c559-44c6-a2f9-0f2a1e876189 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling rejected under rejection sampling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a971f4a7-890b-4a1b-887a-74b13de80b7c · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The TV gradient is bounded byq j
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8158ee0a-c955-4abd-95b2-88fff7cb35bf · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling mode-seeking
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de69c6a3-4937-41a7-8aac-f7c89221e1d3 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The reverse KL imposes anasymmetric 22 penalty: over-estimation (qj >p j, so log(qj/p j)> 0) incurs a much stronger gradient than under- estimation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d381fef1-2418-4f18-baab-034185babcf3 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4f813b-9737-403a-8f47-1ee4b3bef279 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac4e1006-56f6-4b4b-8a67-1cb43cd31550 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Practical considerations.The above analysis assumes δ is a constant, but in practice, the draft head has finite capacity
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0256dc-d046-46a2-85d2-54be6555b236 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling , γ, draw ui ∼Uniform( 0, 1) and accept ˆyi if ui ·q i( ˆyi)<p i( ˆyi), i.e., with probability min 1,p i( ˆyi)/qi( ˆyi)
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d99502d-fb3f-4b77-afa6-dccfd14566c4 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling For rejection at step j, the residual distribution is presid(v)∝max( 0, pj(v)−q j(v)); for the bonus token (all accepted), the residual is simply pγ(v)
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c4527f-6d18-42e6-a7cf-84e12de6c290 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The kernel records the index of the first rejected step
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07634fa9-0f6b-4b60-945c-8a00a5d2d8c4 · outbound
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling For rejection at step j: zresid(v) =log max( 0, pj(v)−q j(v)); for the bonus token: zresid(v) =z target,γ(v) (the raw target logits)
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b39c9c3-e86a-4261-9322-81a945c9f801 · inbound
D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1262a26-beb5-4bf9-b512-c4321cde853f · inbound
AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.