Pith. sign in

Paper Citation Record · LEDGER

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

As of 4 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2606.18967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.18967 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:32:41.428871Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact28
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 401e0ff3-67fc-43ae-8181-4123737467e8 · outbound

This paper cites ShareGPT_Vicuna_unfiltered.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts ShareGPT_Vicuna_unfiltered

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:2ff470dc5da4ad7d12b0d1e90f632060340d2643899651d464102c85078ce769

Observation 74a8250b-186d-43de-9914-e72701a9791c · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:8cfa0a70dddfabc86de9664278e971b84a66bebe082fa575800d8f3ee1c6ca9f

Observation 491c12e2-642f-4243-94bc-ba6578d1f919 · outbound

This paper cites Qwen3-8B_eagle3.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Qwen3-8B_eagle3

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:a31d5055e917fe49dd05cdaa3fb2646c012cbd652794175a94539f4932a8817a

Observation 6b586df9-bb75-4094-b114-bf83e6d953f8 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Lee, Deming Chen, and Tri Dao

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:bf056dbab0fa4fc71f22126fc75b7f743ccf9962f697c5693a8f473919357522

Observation 01a3e45a-ff06-4614-b9ad-7d31f4933a9c · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.225539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:df1a8f7cdac3fb062b099751256b4826a4ddc5e5c55c8b2c562a8efc0509c7c7

Observation d52f6851-713c-4a7b-ae0e-951787e68ebe · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Accelerating Large Language Model Decoding with Speculative Sampling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.244700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:9c530269208bce0b81f6da1b2a952f694cd3f35cb5686b56ac3fc48d1f33cb69

Observation 2d1ea871-053f-44de-ac83-f4fa1cf6ea8a · outbound

This paper cites Clasp: In-context layer skip for self-speculative decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Clasp: In-context layer skip for self-speculative decoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:a2d6e919f929d5b8fd657bde6ce8564a172c40a59356590c944597b71e628c09

Observation 089ab0e4-1fe8-420c-80de-37333be75d7b · outbound

This paper cites Respec: Towards optimizing speculative decoding in reinforcement learning systems.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Respec: Towards optimizing speculative decoding in reinforcement learning systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:34d4122aac2af1290bf86fe26a1d699aab43a225af814e574f4d2daa2a3d273e

Observation f8f1d795-b2f8-405d-946e-684de8676e8b · outbound

This paper cites Do NOT think that much for 2+3=? on the overthinking of long reasoning models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Do NOT think that much for 2+3=? on the overthinking of long reasoning models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:1a86f1e23cd2bd127d5aed51031a9cdb89f735ef203c872117dd795966caf167

Observation f2e6d733-b962-4b59-a5ae-cfb53b81ce57 · outbound

This paper cites Jackpot: Optimal budgeted rejection sampling for extreme actor-policy discrep- ancy.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Jackpot: Optimal budgeted rejection sampling for extreme actor-policy discrep- ancy

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.249152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:0c6db5aecb0393a5d4c89e89c17bc3dac34412efc8c1d41e6a23c47a92b2ad22

Observation 096435c1-5ca0-491d-a364-d197d4d79e39 · outbound

This paper cites Multi-Head Attention: Collaborate Instead of Concatenate.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Multi-Head Attention: Collaborate Instead of Concatenate

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.202274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:471acb148dd288b776c7d2002e3939260e15b957c0bd4388c3448fd50265cc0e

Observation 3a823961-1684-4781-9146-3824217df725 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.282111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:3d517c691b14922fc78f85e7ec5d00ff13581d88e371fd6a4fa5bb3a93841729

Observation df675402-2c1d-418b-839f-2afa0b89b328 · outbound

This paper cites QLoRA: Efficient finetuning of quantized LLMs.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts QLoRA: Efficient finetuning of quantized LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:0a144082ae644f68a5e92fec7517c2462c6228d11d672a076e99bd2851663d58

Observation 61546bbc-5478-4c52-9caa-82cdb170d942 · outbound

This paper cites Marlin: Mixed- precision auto-regressive parallel inference on large language models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Marlin: Mixed- precision auto-regressive parallel inference on large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:97206f6af05bdd3748b11424c079202ac464419f95cdbc31b0b165312f59d559

Observation b815452b-940e-4eda-b942-618666a520a9 · outbound

This paper cites AREAL: A large-scale asynchronous reinforcement learning system for language reasoning.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts AREAL: A large-scale asynchronous reinforcement learning system for language reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:7705e578e5ee7f22ca3c279166b41b44b9d7dd0fe5fd5c4d86922a1dc361aadb

Observation 7b66cd37-9c2f-478f-9b8e-7282c5f24bd6 · outbound

This paper cites Ai and memory wall.IEEE Micro, 44(3):33–39, 2024.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Ai and memory wall.IEEE Micro, 44(3):33–39, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:1582b247d2d253d830b5624c550ac3fb4cb8ecc2f74eeffc7085dd44ac8030cd

Observation b8c02048-5cc8-477a-91c0-3ef9024c4899 · outbound

This paper cites The Llama 3 Herd of Models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts The Llama 3 Herd of Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.245022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:6c1dbe104e44471804e294c1f2a9782da267568ae549e2166d38ec22096fa92d

Observation 5c0f9995-c7ac-4cd2-8fce-2a93d1c4be8f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.248890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:53a6eb17c5ed8184c938e51a3a238737bf2515754f6fd96621f21c0f9d08e1ea

Observation 6ef0ae08-8da4-46f6-abff-172ecf02e0b6 · outbound

This paper cites History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.266551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:337afb33673b0f2ad69f08034606567d88469ef53e3965fd2f08529df41744ef

Observation 03ddd7fb-1475-4523-a9e8-b943733a0d2f · outbound

This paper cites Rest: Retrieval-based speculative decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Rest: Retrieval-based speculative decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:cb466cb0634467a3a31cd5a3650daa8c635e6897b5af3515a2d89fe9475c9a81

Observation df752bd7-8ba6-4da1-b2aa-69fea7e0ced5 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.255776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:49e6d4a201bb9ff7e3a891ce4ef4246a14295f885bb0dfe0105fc45cf49921fa

Observation c28e7810-ae48-440d-9e80-f589b2ddf503 · outbound

This paper cites Taming the long-tail: Efficient reasoning rl training with adaptive drafter.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Taming the long-tail: Efficient reasoning rl training with adaptive drafter

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:4895709639746340e70548a55f02583b27b7b39541dd0ad29ef9ead21a605489

Observation 624c4b42-c59b-4bd6-aca8-64c5d12ac679 · outbound

This paper cites Ash, and Akshay Krishnamurthy.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Ash, and Akshay Krishnamurthy

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:c7b36bc3309707842cbcf4f9026d43310176a65780df7c510c37fcc4be6a446b

Observation b3239e70-32c3-48d7-8522-89eff86a266a · outbound

This paper cites an unresolved cited work.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:d295d07fda3cf07c2af51ae4d87a1c3bd4f6c7cab3824eb38ea1dcbeee3ae14f

Observation e799e847-4719-458c-be4e-93a7cfa8cfac · outbound

This paper cites Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.265004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:553640c1293f11a1a9e4fa48571ac5bd9de35cc3d3fbdd1a853e12a3014b6db5

Observation f579feef-5ee2-4e77-a45a-40e2a210a6a1 · outbound

This paper cites Revisiting Entropy in Reinforcement Learning for Large Reasoning Models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Revisiting Entropy in Reinforcement Learning for Large Reasoning Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.277857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:72d721bd7a4f7b52b5c2c157af22558c303ad7761a88b2566f9baba1af31fe65

Observation 5824601c-62e8-4424-abff-11db2ac62598 · outbound

This paper cites Beyond next-token prediction: A perfor- mance characterization of diffusion versus autoregressive language models.arXiv preprint arXiv:2510.04146, 2025.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Beyond next-token prediction: A perfor- mance characterization of diffusion versus autoregressive language models.arXiv preprint arXiv:2510.04146, 2025

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.269288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:ef6d8200a3c20f0a05112b39f58c4a7be6e6dab1c1778422651046723c4475dd

Observation d9d72b58-3f14-4544-8534-ceba0ab01eca · outbound

This paper cites LLM Post-Training: A Deep Dive into Reasoning Large Language Models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:07.220274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:70e12cfa7cc4837198a412b01f1228330f4ef54f1cbe8b13693ab8ab04676fc6

Observation a0b4c0c4-8816-4b42-a26d-da55a68b3a1f · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Efficient memory management for large language model serving with pagedattention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:118852729c400b844258de752098fb92b53a468bfd9f6b92353aa2cea0c5ef04

Observation 4e25db5c-46f8-452f-afea-d2eccd38ce2f · outbound

This paper cites Fast inference from transformers via speculative decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Fast inference from transformers via speculative decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:e8a58bbfc4c4554344b03f08a3b3047bfa342c5ec649b35e26197453cd510595

Observation 539fca33-6f90-4834-9ec0-911d5121f236 · outbound

This paper cites QuRL: Low-precision reinforcement learning for efficient reasoning.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts QuRL: Low-precision reinforcement learning for efficient reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:b0933a902855a7fd55ba150ac8edf90ec056211083e9d3ddd37ecd555523af9d

Observation 6e84057a-ba8f-49e8-b30f-20eca83479c9 · outbound

This paper cites EAGLE: Speculative sampling requires rethinking feature uncertainty.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts EAGLE: Speculative sampling requires rethinking feature uncertainty

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:604615df1cd5a07a888e9a32ec9e0b1268a2b7171b5ddbc389493309f5eaa083

Observation 689d5d5f-9169-458b-b169-558496cc64a4 · outbound

This paper cites EAGLE-3: Scaling up inference acceleration of large language models via training-time test.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts EAGLE-3: Scaling up inference acceleration of large language models via training-time test

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:bac4519f30a817f1aa02d59fca5a0ea8fdbecfd4cb97a8457d1e8f8f1af14141

Observation 033e901e-b131-4ae1-9a7f-f395d0defdfb · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:4844892718863f3fe44e1249a12573734c387871670fc52484d83bea18357c21

Observation aa27d7c8-80b4-486d-ad77-532f36c3f82a · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.269924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:bb6ead65961b3cc9eab6561ba6e7c01fea9b2356251d8d020bbe7434a00d8a22

Observation 3a33f6f3-f0d3-477b-9a0d-a82932393ea7 · outbound

This paper cites Spec-rl: Accelerating on-policy reinforcement learning with speculative rollouts.arXiv preprint arXiv:2509.23232, 2025a.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Spec-rl: Accelerating on-policy reinforcement learning with speculative rollouts.arXiv preprint arXiv:2509.23232, 2025a

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.273259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:6806030b3dfb2b4fa2e35f0af91fe85583a21b483a10996d78617ddcbef0a050

Observation 002875a9-3515-430f-951e-b0f587047687 · outbound

This paper cites Speculative decoding: Performance or illusion? InNinth Conference on Machine Learning and Systems, 2026.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Speculative decoding: Performance or illusion? InNinth Conference on Machine Learning and Systems, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:c7f19e0d36599c6f5d323f8db2f8a7f44bf28d875777c21160d39e640c383439

Observation e04402bd-41bb-4253-b032-17af098dd397 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Understanding R1-Zero-Like Training: A Critical Perspective

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.273927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:080047e0dc48eba710de13dfc880f0ca33b04067e8456bc823c2281494c722c7

Observation b6d3321f-121a-4e6c-ba11-31d9858187dd · outbound

This paper cites Qwen2.5-7B-Eagle-RL.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Qwen2.5-7B-Eagle-RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:6459940f9e393ae53b0c401dc8f6dbd9ea38b19e7031d1179250a77eb76b6806

Observation 3c7602ad-03e5-4fb9-9c58-1d7b1b9825a4 · outbound

This paper cites OpenThoughts2-1M.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts OpenThoughts2-1M

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:fa217b33cbd71f12d093eb2cd3ce177a1748ee7fcb6a374da4f4c55543ee9582

Observation f42b4ab3-400b-4482-a123-6f799ea2bce6 · outbound

This paper cites Lossless acceleration of large language model via adaptive n-gram parallel decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Lossless acceleration of large language model via adaptive n-gram parallel decoding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:60af71ba2756a0c865bc11c55d8b193d0ca5f4777678aa174cb6fa760aa1d67f

Observation 6c35a01c-f8ba-4f29-9177-06cb39196d6c · outbound

This paper cites Llama-3.1-8B-Instruct-speculator.eagle3.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Llama-3.1-8B-Instruct-speculator.eagle3

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:4ce663ecdd92085feea7922674e53f064cb6a917baa1eb2ee6829d18b6213a02

Observation aed89340-b385-42cd-9ae4-0bb4e88f518a · outbound

This paper cites Qwen3-8B-speculator.eagle3.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Qwen3-8B-speculator.eagle3

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:404bbf8419d39e9e5418d55e37a1ca15d9c7133dbdfc8604e5e80968ee7a0366

Observation ec8d9786-dc88-43df-8f96-87e013c0998b · outbound

This paper cites Qwen3-8B-Thinking-speculator.eagle3.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Qwen3-8B-Thinking-speculator.eagle3

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:ddc3e2dd8d5eeaac00597253596426981670191283e01b2f28496dfa25d6856d

Observation 07ae8e72-244d-4f0c-a86e-79ab460cecc1 · outbound

This paper cites Magicdec: Breaking the latency- throughput tradeoff for long context generation with speculative decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Magicdec: Breaking the latency- throughput tradeoff for long context generation with speculative decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:e065ce40e170ee044ec219f099ea3c67f602570d5f7e702aaa52b9dccfb6fb1d

Observation 5162ea17-7fec-4162-9a11-328b89e201be · outbound

This paper cites Proximal Policy Optimization Algorithms.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Proximal Policy Optimization Algorithms

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.196426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:25bec777873892c2fd121bd32a08254b4bdfa7591796c0416e53077df1987a47

Observation 5eae0e86-037b-41b2-8f18-04404ee066ed · outbound

This paper cites Beat the long tail: Distribution-aware speculative decoding for RL training.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Beat the long tail: Distribution-aware speculative decoding for RL training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:181f90b1935d83ede7d6b4ef169e52f283788049632ad0a09f7c0631e701fdfe

Observation d525bc95-f198-4347-a670-cfe530223610 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.192273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:10955bdf84e3486e4dde533e0a22a75a14036159ea25f7367cf50838a0419cf6

Observation 74a57d62-c63a-4c76-927f-0ebce3fb9472 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Fast Transformer Decoding: One Write-Head is All You Need

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.200433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:9549a0e457f09e729339c12e2fae8db172cabf5873671836be8d01c104109a2d

Observation ad3e1935-6b89-458a-97bd-c92e83aee2d0 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Hybridflow: A flexible and efficient rlhf framework

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:32de3fe737030ac4854b2e2ac27701409399887c5f847edf3a0a743a3904db3f

Observation 4a998feb-e360-4f50-abbd-524b958259e9 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.221087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:794d3852067198ba595e8fb52b84a18c8756eaca031a1fadc41ee813131a8337

Observation 2c0eaa85-2e43-4fde-a8f1-1c490703b6b3 · outbound

This paper cites Knn-ssd: Enabling dynamic self-speculative decoding via nearest neighbor layer set optimization.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Knn-ssd: Enabling dynamic self-speculative decoding via nearest neighbor layer set optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:66985c1485e262f85a65ed7e82c0b734e930b775f9b84fba6f798b0aa8a45a33

Observation 3b1b71bf-0513-4e60-9368-e20f9ecc045c · outbound

This paper cites The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.188814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:0ed149caab46be4fcf97085428a289e2f420e8809b34e5d307966b28f8cb739b

Observation 591e8ae2-7f7b-4d52-99d0-b2491c32bbd1 · outbound

This paper cites QUEST: Query-aware sparsity for efficient long-context LLM inference.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts QUEST: Query-aware sparsity for efficient long-context LLM inference

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:b5f14d2c6a33a2adee0f765e861ed2a501205723c2e60b6a41d26363635ade11

Observation 775e39a8-a16a-4cb1-8da9-ad2f6ad83792 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Kimi K2: Open Agentic Intelligence

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.217202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:0cc4e820c048905947f003a3851277d16fab6d9e88bbc110fa5d38f2190aaa1b

Observation 5e44f48a-6cfc-48c4-90a0-fed1cf1991bf · outbound

This paper cites Mahoney, Kurt Keutzer, and Amir Gholami.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Mahoney, Kurt Keutzer, and Amir Gholami

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:89544c1895d9135e874fe2927a0331ec40a8021f17f393aa43331b1ce78888db

Observation 6bcbeb72-5cf0-4aeb-9c79-2e56942c88b5 · outbound

This paper cites Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:80aafc92b3c1580690d4e4a49df31b927eaa28ba51b68d3ba44340b6044d1635

Observation 5570306a-6e39-4318-be6d-bfd10c89c666 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.Communications of the ACM, 52(4):65–76, 2009.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Roofline: an insightful visual performance model for multicore architectures.Communications of the ACM, 52(4):65–76, 2009

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:aef6f6c31b16e28edd5f24d14bad7123f33c5e13d2ae99ff6d66f8ba10ed0b41

Observation 2c9f3bfe-dd90-4519-9d1b-3385a3b5e572 · outbound

This paper cites an unresolved cited work.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:e2a7318b28ce2a3e1c4ed44a4a7e0a9c7bcd4132f5373dd0dce1fa28ac7000b6

Observation 3f1fccc6-92c5-41b5-8ed2-9327557dccdf · outbound

This paper cites SWIFT: On-the-fly self- speculative decoding for LLM inference acceleration.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts SWIFT: On-the-fly self- speculative decoding for LLM inference acceleration

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:8101efac52ce563e5df36fce52e5f43f525f96ad56ae07391e6fd8c1a2da6ed8

Observation 22659c0b-b384-40a3-a05f-660fe3ec4b88 · outbound

This paper cites Qwen3 Technical Report.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Qwen3 Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.286798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:8a0c3faeb020cd4a54610f5ad4ed3cc6ab677ec5365b0bbc064002b3ce874769

Observation 3f83e017-9b04-43e7-8638-b5e75446d455 · outbound

This paper cites LLM Probability Concentration: How Alignment Shrinks the Generative Horizon.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.260894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:e93cf4b46a2321f0ca49a4f1157257d70b8b0dbf87b8ab67ea7c25bb8c885243

Observation 929eabc7-458a-4c90-9fdf-ffc2ba7be07e · outbound

This paper cites Qwen2.5 Technical Report.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Qwen2.5 Technical Report

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.204186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:0d41cb7ff19e0eb34b4f0bb9d9b85f8d016b472ce6443fe4c1e2bc76cdfb84c2

Observation 32bc576d-d5fc-436c-afa3-67ba8ac95c74 · outbound

This paper cites DAPO: An open-source LLM reinforcement learning system at scale.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts DAPO: An open-source LLM reinforcement learning system at scale

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:cc1eb0e887d2eb7be161ba1f88bcf5374c6025d6840617cc524cdc3a18af0b3c

Observation d58a23c3-f7e6-4a53-997d-a141d21b67c9 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.177818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:ec76813f526fe496cf3622a040a85dd0b40572d994a4e29ef678b7ec33314e4c

Observation 41f815be-0c80-48af-91d9-4f25852e9020 · outbound

This paper cites Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.186895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:7f14c8240c28128946c535dd89960ae6a8f57486b00cc3a0ace32895d1e9fdd6

Observation 9f7c7221-9750-4376-9a02-c799102237f1 · outbound

This paper cites EAGLE3-LLaMA3.1-Instruct-8B.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts EAGLE3-LLaMA3.1-Instruct-8B

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:4cb91e64a638fbec86aa739340a854d26fe1812c33e2d23875892ba3579463b0

Observation 63714f56-12bb-4c06-a05f-6d7ee12f4b93 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:07.290498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:9a4b084fe207c512d76e060a030dec9ceae3514a25638f2fb59aaddb30f8de3a

Observation 273489d2-45c8-436a-86ad-a3e7d637df99 · outbound

This paper cites SimpleRL-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts SimpleRL-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:2131032f3249ddb839572aa2574f521fe52169d83a2f27f1b1d439851f2bdd2f

Observation 44ba7741-56cd-46cb-8991-406ebcc61ac3 · outbound

This paper cites Draft& verify: Lossless large language model acceleration via self-speculative decoding.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Draft& verify: Lossless large language model acceleration via self-speculative decoding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:98b63764bab583158f55bef2972bf3dd482d46e100907454ac25370aa67dc279

Observation 38ab0e32-9700-4cec-8cd9-fae29696699b · outbound

This paper cites Sortedrl: Accelerating rl training for llms through online length-aware scheduling.arXiv preprint arXiv:2603.23414, 2026.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Sortedrl: Accelerating rl training for llms through online length-aware scheduling.arXiv preprint arXiv:2603.23414, 2026

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.237657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:6629f1d84cb4757ab99c93c318a8559c9fc0ac4e17c924c71bfdca9ef7ebdfb0

Observation 6acab2f0-2517-49cb-9bd5-a122ac62e56c · outbound

This paper cites FastGRPO: Accelerating policy optimization via concurrency-aware speculative decoding and online draft learning.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts FastGRPO: Accelerating policy optimization via concurrency-aware speculative decoding and online draft learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:0ec30793ea39f7575219c604aa2bd36b67d80a095248e0567a4891c7bdfb5ea7

Observation 1808f50c-6a46-4344-9807-5866e98943b9 · outbound

This paper cites Kakade, Cengiz Pehlevan, Samy Jelassi, and Eran Malach.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Kakade, Cengiz Pehlevan, Samy Jelassi, and Eran Malach

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:40215a70904b8d451d84bf4c9abce421b4214b54ba35f288d6b8916183727a84

Observation 0834e3f0-9eed-415a-a7db-9437ded4256b · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Gonzalez, Clark Barrett, and Ying Sheng

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:4fefcd018d8211b09ad078b9bdc4e0f75b4bc74d138686122dd385b5a4ab382a

Observation f4e6a1f5-29bd-417a-94f0-b55b764beccb · outbound

This paper cites Distillspec: Improving speculative decoding via knowledge distillation.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts Distillspec: Improving speculative decoding via knowledge distillation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-26T21:32:41.428871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:84c1692cf1af80f5e268c61fa1b086aff107b319c7803e43daaf60a60a4af93b

Observation 37327a86-a726-488a-8257-ecb4590938ee · outbound

This paper cites April: Active partial rollouts in reinforcement learning to tame long-tail generation.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts April: Active partial rollouts in reinforcement learning to tame long-tail generation

Reference 76

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T23:59:07.241728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:ab4440b7f5e6f4c76a376982017ba27805f41f96de01bd6165714fe22b0b3d23

Pith citing papers

No inbound Pith citation observations are available.