Pith. sign in

Paper Citation Record · LEDGER

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

As of 7 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 2 inbound Pith citation observations for arXiv:2606.12370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.12370 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:34:48.709434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a1fdafe-eeae-436e-867c-e1b7bb751525 · outbound

This paper cites Draft-OPD: On-Policy Distillation for Speculative Draft Models.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Draft-OPD: On-Policy Distillation for Speculative Draft Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-27T10:30:51.814090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2d6c7436ef46e1db986c12c5843535239e72d45dee2e9777e2fc155eb01a0218

Observation 5d0b5a34-4160-4c2e-ad9a-8c1252c72cc3 · outbound

This paper cites The TV gradient is proportional to qj, so it automatically ignores low-probability tokens.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The TV gradient is proportional to qj, so it automatically ignores low-probability tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:104a0fb498ffacc586a45d9fc432e3d2b0c99592dbb2c2387ad25e62f6e5c60b

Observation 358997f3-c559-44c6-a2f9-0f2a1e876189 · outbound

This paper cites rejected under rejection sampling.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling rejected under rejection sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:af77dd59e75db4b277917c68f31742486b769ac5ea58e38509596c0fe7a736ba

Observation a971f4a7-890b-4a1b-887a-74b13de80b7c · outbound

This paper cites The TV gradient is bounded byq j.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The TV gradient is bounded byq j

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:7b7d602cb80c28cefd7331fda96e96f4c414e32db6827b5d1513688ff74d79b5

Observation 8158ee0a-c955-4abd-95b2-88fff7cb35bf · outbound

This paper cites mode-seeking.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling mode-seeking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:14c34a74a78080f066fe82e4ac70454321f8f0a4dc09ba640e9d05d7cfa7a225

Observation de69c6a3-4937-41a7-8aac-f7c89221e1d3 · outbound

This paper cites The reverse KL imposes anasymmetric 22 penalty: over-estimation (qj >p j, so log(qj/p j)> 0) incurs a much stronger gradient than under- estimation.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The reverse KL imposes anasymmetric 22 penalty: over-estimation (qj >p j, so log(qj/p j)> 0) incurs a much stronger gradient than under- estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:15e1a050797d479dfa1636c7f7c3b880b8e134ccb58a80224a78129c8216f8f6

Observation d381fef1-2418-4f18-baab-034185babcf3 · outbound

This paper cites an unresolved cited work.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:538e780de18e1a405bf665a899219763193b83c580a2fd23e637225fdb70d509

Observation 3f4f813b-9737-403a-8f47-1ee4b3bef279 · outbound

This paper cites an unresolved cited work.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:96896c65ef38c84ec42755058bce9c595cf7ebb93122233c7741ba160d05917b

Observation ac4e1006-56f6-4b4b-8a67-1cb43cd31550 · outbound

This paper cites Practical considerations.The above analysis assumes δ is a constant, but in practice, the draft head has finite capacity.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Practical considerations.The above analysis assumes δ is a constant, but in practice, the draft head has finite capacity

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:1f760f7889fd84e1c0a22bfc385b558f68c79bdcf8b47441d356d69694f1f582

Observation bf0256dc-d046-46a2-85d2-54be6555b236 · outbound

This paper cites , γ, draw ui ∼Uniform( 0, 1) and accept ˆyi if ui ·q i( ˆyi)<p i( ˆyi), i.e., with probability min 1,p i( ˆyi)/qi( ˆyi).

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling , γ, draw ui ∼Uniform( 0, 1) and accept ˆyi if ui ·q i( ˆyi)<p i( ˆyi), i.e., with probability min 1,p i( ˆyi)/qi( ˆyi)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:f4fdb01af6c67ac85e3144d555831676dde79ff68572d4fde414aeb1478256ff

Observation 8d99502d-fb3f-4b77-afa6-dccfd14566c4 · outbound

This paper cites For rejection at step j, the residual distribution is presid(v)∝max( 0, pj(v)−q j(v)); for the bonus token (all accepted), the residual is simply pγ(v).

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling For rejection at step j, the residual distribution is presid(v)∝max( 0, pj(v)−q j(v)); for the bonus token (all accepted), the residual is simply pγ(v)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:e57920c093dbdda809170802f42a9cfe961a07d4e41d0382d8cb0478143935ce

Observation 18c4527f-6d18-42e6-a7cf-84e12de6c290 · outbound

This paper cites The kernel records the index of the first rejected step.

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling The kernel records the index of the first rejected step

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:e8bab4701303f326a3b3aa0d6c0bf5a289c759501195418e293d8805248685d6

Observation 07634fa9-0f6b-4b60-945c-8a00a5d2d8c4 · outbound

This paper cites For rejection at step j: zresid(v) =log max( 0, pj(v)−q j(v)); for the bonus token: zresid(v) =z target,γ(v) (the raw target logits).

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling For rejection at step j: zresid(v) =log max( 0, pj(v)−q j(v)); for the bonus token: zresid(v) =z target,γ(v) (the raw target logits)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T10:24:34.257293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:24:34.257293Z digest=sha256:2c2177a6aae07f07e63050ae19fd7880f25123fe24ad77d7f147e1bdf4118bee

Pith citing papers

Observation 7b39c9c3-e86a-4261-9322-81a945c9f801 · inbound

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding cites this paper.

D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:48.709434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:34:48.709434Z digest=sha256:a89d6a1844eddeaf59fdf9712ca7d826e5159312f0084c2e9a82f220354dce4d

Observation b1262a26-beb5-4bf9-b512-c4321cde853f · inbound

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding cites this paper.

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T01:22:13.804729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:22:13.804729Z digest=sha256:7e22dd1557da6d17f42a61852f98fb11fa849a6664a7af70fc7ffb8a20a45bbe