Pith. sign in

Paper Citation Record · LEDGER

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 3 inbound Pith citation observations for arXiv:2606.12370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.12370 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:43:07.502284Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T14:43:07.950091Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a1fdafe-eeae-436e-867c-e1b7bb751525 · outbound

This paper cites Draft-OPD: On-Policy Distillation for Speculative Draft Models.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Draft-OPD: On-Policy Distillation for Speculative Draft Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-27T10:30:51.814090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:dcbbddd236d4f625b2c5e28574cd156b408f80eeedae67bce99a8b21ea57cef2

Observation 5d0b5a34-4160-4c2e-ad9a-8c1252c72cc3 · outbound

This paper cites The TV gradient is proportional to qj, so it automatically ignores low-probability tokens.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The TV gradient is proportional to qj, so it automatically ignores low-probability tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:b0d3b6d58432b500c51d0979600a7786a00b8cc5bd149d5041f98b46808a4de7

Observation 358997f3-c559-44c6-a2f9-0f2a1e876189 · outbound

This paper cites rejected under rejection sampling.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling rejected under rejection sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:680ccdcd42d88f2cfc9ea0e1d25e40f816e3fd332731cc22bce02ba1535120e9

Observation a971f4a7-890b-4a1b-887a-74b13de80b7c · outbound

This paper cites The TV gradient is bounded byq j.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The TV gradient is bounded byq j

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:b1db78944c750f7352bd680656f3c40a21a9395c16aa4ef3738431fcca79c4ba

Observation 8158ee0a-c955-4abd-95b2-88fff7cb35bf · outbound

This paper cites mode-seeking.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling mode-seeking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:5b4159dc29057b2e1bdfd637948241b14b332fd78c7f9837f953ec8b2155148a

Observation de69c6a3-4937-41a7-8aac-f7c89221e1d3 · outbound

This paper cites The reverse KL imposes anasymmetric 22 penalty: over-estimation (qj >p j, so log(qj/p j)> 0) incurs a much stronger gradient than under- estimation.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The reverse KL imposes anasymmetric 22 penalty: over-estimation (qj >p j, so log(qj/p j)> 0) incurs a much stronger gradient than under- estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:42a9882cc22b4d463c4a54f1ee743533c5cc8486c85ec48df03118be7617a3ed

Observation d381fef1-2418-4f18-baab-034185babcf3 · outbound

This paper cites an unresolved cited work.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:92cd1b9ecefc6e054d225d071d418f7c8d639b4a64803b6d35865d00d61eb08c

Observation 3f4f813b-9737-403a-8f47-1ee4b3bef279 · outbound

This paper cites an unresolved cited work.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:e7bcdfe6f1d71640b6e418ac7e429812642adec2ca57764bbcb45ae8140a2f29

Observation ac4e1006-56f6-4b4b-8a67-1cb43cd31550 · outbound

This paper cites Practical considerations.The above analysis assumes δ is a constant, but in practice, the draft head has finite capacity.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Practical considerations.The above analysis assumes δ is a constant, but in practice, the draft head has finite capacity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:0f0cfe8a7e48ac6e85a0670962a453244b38adac15543bf19b9b996daa1c5be2

Observation bf0256dc-d046-46a2-85d2-54be6555b236 · outbound

This paper cites , γ, draw ui ∼Uniform( 0, 1) and accept ˆyi if ui ·q i( ˆyi)<p i( ˆyi), i.e., with probability min 1,p i( ˆyi)/qi( ˆyi).

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling , γ, draw ui ∼Uniform( 0, 1) and accept ˆyi if ui ·q i( ˆyi)<p i( ˆyi), i.e., with probability min 1,p i( ˆyi)/qi( ˆyi)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:005695075a8e91ab5c6588c2488ff38fd466c35e6cecfb3499f6b22b0b537754

Observation 8d99502d-fb3f-4b77-afa6-dccfd14566c4 · outbound

This paper cites For rejection at step j, the residual distribution is presid(v)∝max( 0, pj(v)−q j(v)); for the bonus token (all accepted), the residual is simply pγ(v).

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling For rejection at step j, the residual distribution is presid(v)∝max( 0, pj(v)−q j(v)); for the bonus token (all accepted), the residual is simply pγ(v)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:ffc3b082954bc18636cb213c9eb1978dca5db71af76d16434ae0cbc13f71695e

Observation 18c4527f-6d18-42e6-a7cf-84e12de6c290 · outbound

This paper cites The kernel records the index of the first rejected step.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The kernel records the index of the first rejected step

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:086f2951b542596f2b314182c20527da3f16593d5495b2c9a19ec565c7005a27

Observation 07634fa9-0f6b-4b60-945c-8a00a5d2d8c4 · outbound

This paper cites For rejection at step j: zresid(v) =log max( 0, pj(v)−q j(v)); for the bonus token: zresid(v) =z target,γ(v) (the raw target logits).

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling For rejection at step j: zresid(v) =log max( 0, pj(v)−q j(v)); for the bonus token: zresid(v) =z target,γ(v) (the raw target logits)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:270f26c26220a796d3ee871cfbb3f23fb3390f312c81c625ba6095fcc9eafe03

Pith citing papers

Observation 7b39c9c3-e86a-4261-9322-81a945c9f801 · inbound

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding cites this paper.

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:48.709434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:48.709434Z digest=sha256:0f52b1a0a2c7a4f79140d708a0a2d3ecc7d376fa1620033411efbde932c3c6ce

Observation b1262a26-beb5-4bf9-b512-c4321cde853f · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:13.804729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:13.804729Z digest=sha256:8705921971d9307641e88d0357ed5df023e3641565a8e25a3e4c27cb36763672

Observation 29174369-95b2-41da-8d45-7dd7c93f3d9b · inbound

K-EXAONE 2.0 Technical Report cites this paper.

K-EXAONE 2.0 Technical Report Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:43:07.957244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:43:07.502284Z digest=sha256:f7f5d67216ce198f93ff3fca84af3ec7d3a576c2e008470f5816e702d853c264