Pith. sign in

Paper Citation Record · LEDGER

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2605.08862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08862 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:23:13.071107Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact20
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e05570b8-bc9a-4e10-a03d-c3e6e8b59ed5 · outbound

This paper cites Polaris: A post- training recipe for scaling reinforcement learning on ad- vanced reasoning models, 2025.URL https://hkunlp.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Polaris: A post- training recipe for scaling reinforcement learning on ad- vanced reasoning models, 2025.URL https://hkunlp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T07:24:49.626886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:9896c5e5bd48a5fa252554b54692677d6cdaeb436f16fc7e12fa5f47f345a813

Observation 5100b6ed-3547-4220-9fd4-f64ec2cdb5be · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.715459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:48bdb27d93e316a6c5c6927ceb9f2cf4f8a23e62393d7f2c7a70f82bb9eb7372

Observation 47ed2012-08f6-4c70-8cce-43db2b63052c · outbound

This paper cites Qwen2.5-VL Technical Report.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:30.903338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:d2551b026a58bb1e153255dbeff2ec2b9d57ca5d3dc507dbdcb9ca6060e7e353

Observation fe2067c4-2c78-4ae1-81d8-15965985ecca · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.892007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:85331319d6a4b18223143abcb340750ba952f9906d7eca75c28f7772044e473d

Observation 8d917e05-0452-4a23-8e23-8177d57a79de · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:4c6e6a18147fcf72d98a21b97533873ef904627eb5c8f287588251d6eab991b2

Observation 32b95ca7-e9dc-4256-933a-b5ab4e5aecab · outbound

This paper cites Soft Adaptive Policy Optimization.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Soft Adaptive Policy Optimization

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:14:32.897338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:9067e907e4d1d974ee1a22c78ec861da0d85dad46dae04f8dbfa2021f08722f2

Observation ea386b91-1233-496f-a36f-0bf3911ab6ef · outbound

This paper cites History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.843384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:90918c793c068554589a9185c0ee2d19500bd1061c780de20df179ca52eafb76

Observation f1b05fe5-3698-47e2-8fc0-1362b22e2cac · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:28:57.414349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:ecfc847c6c706362064ec791d772ce24f2f6f2ad2ac1f2cc5bbb1515e6549a01

Observation af40d67d-8f32-42b6-8d77-c2325c8385e9 · outbound

This paper cites Taming the long-tail: Efficient reasoning rl training with adaptive drafter.arXiv preprint arXiv:2511.16665.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Taming the long-tail: Efficient reasoning rl training with adaptive drafter.arXiv preprint arXiv:2511.16665

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.894242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:9869e1ad3f0ba757123d3808a5d9925223a433f6162081972106c327750a7f17

Observation 6a00875e-4e0d-479b-a8a8-3abf1ebd103a · outbound

This paper cites OpenAI o1 System Card.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning OpenAI o1 System Card

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:30.787686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:b75d8226de54d1b0afa8592a8ef9adc199c9a0019eebb74d207ce22fcd2d227d

Observation 3f2f4a67-2f64-4c34-8b46-44c4e5e3ebbb · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:2485ad59b038504755e287f3b07591a33ae6669309f287993c6f6292be08c5b6

Observation 914c09ce-b2cc-4343-875f-9fa281a1dc1e · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:15:49.524886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:244ec250ffb747643811cfaf143d5d1193f5011b3832b02288af85670f9fbb56

Observation 0d7ffdfa-24d6-45e9-ac6b-83cbffc4dd0f · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:52:19.804345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:0e82993408cf85b9ef57129f53722bdf595765f220634143716b4b73d4835cb5

Observation 38067cc4-94fc-4f5a-a2ee-f5eda571a40a · outbound

This paper cites Spec-rl: Accelerating on-policy reinforcement learning with speculative rollouts.arXiv preprint arXiv:2509.23232, 2025a.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Spec-rl: Accelerating on-policy reinforcement learning with speculative rollouts.arXiv preprint arXiv:2509.23232, 2025a

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.959383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:ca6653af27496258430e2cc55cf390eda1a595ac97acf56ab4ac3007f9cf0d39

Observation ec2eace2-6a9f-4f50-bb0c-25f4a22b274b · outbound

This paper cites Faithful chain- of-thought reasoning.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Faithful chain- of-thought reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T07:24:49.631190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:eaba750b26132b5dbdb080684aa38d6f89f6461a4f7922112774ff9aa44d79a9

Observation ed7516dc-0572-481a-b551-4871e40c5470 · outbound

This paper cites Defeating the training-inference mismatch via fp16.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Defeating the training-inference mismatch via fp16

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.993342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:194892d963905d6686e1121a792a144641c44a5798a73c8b051b4ae3c67dcce5

Observation 5ba84f2b-f5dc-457e-aafc-24365278e0ca · outbound

This paper cites Proximal Policy Optimization Algorithms.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:01:30.924069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:3e0e9d42bca044f9e30567874f0d63cdc02ddd108bb837c59d319ab7652d67f0

Observation 9c561d49-c98d-4e75-9a09-5e0ab25d70db · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:30.936203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:05dbe62c7507480af061ab8573c53b843df7e884310a99c12bec0166833c95a8

Observation 711743ff-97a3-423d-aea1-7aa6529bb2a3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:30.969467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:d39ec190a7ccbbb50435a0ce0c0f967032d02fffb5f44b8ffd366c1ce07c66d8

Observation a1fd7f19-e681-4c2e-b9c0-38c35549ed28 · outbound

This paper cites Jet-RL: Enabling on-policy FP8 reinforcement learning with unified training and rollout precision flow.arXiv preprint arXiv:2601.14243.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Jet-RL: Enabling on-policy FP8 reinforcement learning with unified training and rollout precision flow.arXiv preprint arXiv:2601.14243

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.749988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:4edb4085bad670c4c137bf309a912e55b7074f912997ffaf73184b6171cf688b

Observation 95ecc842-2b36-4188-b412-707249b926f4 · outbound

This paper cites Jones, Zheng Dong, and Peipei Zhou.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Jones, Zheng Dong, and Peipei Zhou

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:26:16.160665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:8105e6f3d1dcc553876b85f610df712c3c55a88b0290393b888fea00ac57a4dd

Observation 7cc00459-56c1-4fde-b935-6ce7b469bc7e · outbound

This paper cites Gen- eration meets verification: Accelerating large language model inference with smart parallel auto-correct decod- ing.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Gen- eration meets verification: Accelerating large language model inference with smart parallel auto-correct decod- ing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T07:24:49.635744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:c7ab2371534b852c5481e6cc32ce2862f40b3af61ca9ac8672a96074b973601b

Observation b09292b2-e795-4a3b-aedc-887f4605f136 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:30.740386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:aa3c4893d37792d26706cd0e2eb4a9a9bc87d8838ea6c41e70117d8d5a598476

Observation f6446d80-5e02-4200-b38d-980f0fa53438 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:21:05.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:ba4ae7ead777702234dcf1782cd28eff3112cb62f48497870fea29ed5ef32326

Observation f10b9858-4408-424d-b838-35a2ac468996 · outbound

This paper cites Group Sequence Policy Optimization.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Group Sequence Policy Optimization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:30.728922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:90251a80b6917c16c38a3698eadf13ff9217844317c1fd9de426a1bd2b1f63a0

Observation 2b31c08b-f0c0-41bb-b1a6-12f9e9f58af4 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:30.777660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:f78e2c41e68abe43431bccce49e0f2d1d07e2499e58818a0259a0e255acb84cd

Observation fc0adcb2-c068-4a48-9360-207c7538cf3d · outbound

This paper cites an unresolved cited work.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-14T07:24:49.622140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:3e979360925ad4c221ad4ece20ebafbb7c8a410f65dc69396d9de3312f80a1eb

Pith citing papers

No inbound Pith citation observations are available.