Pith. sign in

Paper Citation Record · LEDGER

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.01651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01651 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:56.846403Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6329b19f-2260-42fd-b9bf-2ab9f8c1b84a · outbound

This paper cites Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.124106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.695540Z digest=sha256:49eeb86cb68242251c534c5ff6336ad1f8f85e9b1c5b993dfe3db6f894972ecf

Observation 9c3e815c-d5ac-46f1-91fd-671471c8f9f9 · outbound

This paper cites Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-04T23:34:57.854572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.698868Z digest=sha256:2de07971e3bbf0c46185d897ec41b38310bf61349560395241b0a6a38d3c0526

Observation b358d3a9-81ac-414e-9fe2-f908aaa5f5a0 · outbound

This paper cites Program Synthesis with Large Language Models.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.701873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.701873Z digest=sha256:b11f9f005cacac17ab82beaad8be7d8daddec3926118bc66d3b146763946db44

Observation a2d1821c-8c83-48db-91bd-4d22083b50dd · outbound

This paper cites Medusa: Simple LLM inference acceleration framework with multiple decoding heads,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Medusa: Simple LLM inference acceleration framework with multiple decoding heads,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.116183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.704970Z digest=sha256:602b94a053489d43dcdc1a5ef9accdfb656c327873329c3b2a4ff6970d66dae1

Observation 5d56c3e1-8fa1-4da9-a50a-4006741a1276 · outbound

This paper cites Accelerating large language model decoding with speculative sampling,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating large language model decoding with speculative sampling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.108000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.708145Z digest=sha256:2440a9fc5266375d49a80107f2afc068a0eb099eaf8a6a1d96ba8cf02bee3df8

Observation 41c3e34a-2038-4378-b496-0fa3be60e862 · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.714038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.714038Z digest=sha256:3085987e8f2ec494b140299b37eed0e342613ba6a9adab08b765c2cb556fdfb7

Observation 3a81927f-d5c7-4993-a494-1ce773e71ff2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.717585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.717585Z digest=sha256:e4d0de41b4dd9fa1073d267c6de07d6a51a351797bc20592e23ac2a5f761717f

Observation 296a7cc9-540f-4ea6-8f4a-de259765a09a · outbound

This paper cites Better & faster large language models via multi-token prediction,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Better & faster large language models via multi-token prediction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.099691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.720709Z digest=sha256:9322dcb7888a3823675b2e0907e2e1eba1a26f05bdeac6503d09b2536728f4c5

Observation f03a369e-2f3c-4956-b816-dc933ee702c4 · outbound

This paper cites Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.092451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.723608Z digest=sha256:33645ce85cbd8a0d2963e570f498533111950eec0f7e6124ebf63abe52ed1993

Observation 94009cd4-5c80-4777-aab6-25484e9187ca · outbound

This paper cites Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.726380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.726380Z digest=sha256:b7ef439f47595941edd57897e01b49bf4c576dff5bc4d9f9fdf8b89383c08c0b

Observation d65ce328-f1ed-48ce-9452-d1ba3ef61784 · outbound

This paper cites Bridging draft policy misalignment: Group tree optimization for speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Bridging draft policy misalignment: Group tree optimization for speculative decoding,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.084724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.729247Z digest=sha256:6e0810f79bc1ecc34d21359ac609aeaf0155462c90fe45b8571f43a5c142d298

Observation fa33ab29-b983-483f-aaff-9bd1176265f0 · outbound

This paper cites Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.734525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.734525Z digest=sha256:c86fdf951d9c1054721d7f5a584492037fc6cfff520300cadca76f1d6a67a3e5

Observation 819522c8-b345-4da7-b4b4-19d448e8f0f9 · outbound

This paper cites Pimba: A processing-in-memory acceleration for post-transformer large language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pimba: A processing-in-memory acceleration for post-transformer large language model serving,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.737160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.737160Z digest=sha256:a65a4e3a6fdfdca96b8ab8e1732249f11a7f10be242ec3ec820bd166c8b8f457

Observation 4b98fde6-42c6-4d26-b3d5-2e0de4616326 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient memory management for large language model serving with pagedattention,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.739878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.739878Z digest=sha256:0caed4adc9a4d5d94ec76fbe47792c07ed626032d50fe9a691a7ee2b2032a54b

Observation a3a41fce-23c2-4e55-a210-5bfde570bdf1 · outbound

This paper cites Fast inference from transformers via speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Fast inference from transformers via speculative decoding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.076505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.742647Z digest=sha256:0b75cab9912ea7ac197c5dbb94bcaad8aef9f449e341c45d499d22fc5af1c616

Observation 510c84a3-2a10-4a5c-8067-64be9b2152fd · outbound

This paper cites Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.748666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.748666Z digest=sha256:d8d3ec38a58a157f1a40dc78bf06683055ee1be0f16b945565f9cc1b5b9f3b43

Observation 8490ba93-2d01-485a-a550-cf3627768e51 · outbound

This paper cites EAGLE-2: Faster inference of language models with dynamic draft trees,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models EAGLE-2: Faster inference of language models with dynamic draft trees,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.068928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.751295Z digest=sha256:9ae8c48b4e490726ce2c376e43294071632b4e8a17ffeee56ff00a6fae2a5281

Observation 343c301c-3e3d-45d1-80b2-3e1d6974f9be · outbound

This paper cites Eagle-3: Scaling up inference acceleration of large language models via training-time test,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Eagle-3: Scaling up inference acceleration of large language models via training-time test,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.061487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.753852Z digest=sha256:f51a926a80a97608c9227ecf4b4b4be8aa3c4251e838a7aec8fc37dfe7744704

Observation 7b0a1622-72f0-44b1-8409-f9b8015de48a · outbound

This paper cites Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.756424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.756424Z digest=sha256:69ad1f54e002fd4ce761a6160c4c6102bf2d90d4d45c8388056792d3239031d0

Observation 541cab22-df40-40a4-9be6-1db16051545a · outbound

This paper cites Speculative decoding: Performance or illusion?.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Speculative decoding: Performance or illusion?

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.052560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.759847Z digest=sha256:3931ac7872fc01c7d18e80e389e745e7077e7017ba6e1fdc8788624eec0cf43c

Observation ef9c02c0-d2bf-4737-bb93-dbf8600041bc · outbound

This paper cites CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.045015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.762499Z digest=sha256:174b2b68b5e5c0f8a401ec4c148d1c79bfa603d132ff044d8cf0206626a3003a

Observation 3460b5a7-58fa-4aa3-ad95-8654d9c6bff3 · outbound

This paper cites Agentix: An efficient serving engine for LLM agents as general programs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Agentix: An efficient serving engine for LLM agents as general programs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.036706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.765160Z digest=sha256:c177ff1f5c2c9a9aee65a294a3cc11cd0565cacdec97ba1f85a0ab4307f6f396

Observation 7650ec64-9a17-41cc-887e-007fdee0c498 · outbound

This paper cites No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.029402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.767778Z digest=sha256:3987285f41f5cf26e2b8fd08615994bd0e6ca0f1cbcfe6750b09b853ec07301d

Observation c00efa0f-bb4c-44b4-a9ea-f5bb1b31479f · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.770410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.770410Z digest=sha256:77a2f2c8d5eccd357b54390733bc2891ccf001f5a117d99b89ed07b4a4c509f7

Observation 3102521f-27a9-443d-aa13-b1c30c48ac49 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.773049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.773049Z digest=sha256:24e1688d72433a8751e60215bc355ddcd1fa78155d3e9f68290e093a550ee042

Observation 1cbac254-604b-47ce-bfb8-e22dd2af5948 · outbound

This paper cites CUDA Programming Guide: CUDA Graphs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CUDA Programming Guide: CUDA Graphs,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.020532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.775582Z digest=sha256:db662508c35d2bf392ed6b48496549e6a132e2ba31b671b55a3bb24f2118a444

Observation 4c92c93e-63c6-4bb4-8c3f-29821dd13de7 · outbound

This paper cites NVIDIA Nsight Compute Documentation,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Compute Documentation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.012710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.778082Z digest=sha256:76263beba1822513a76c1e84974c460d0c45fc4bc30f9739fb48b233f0c8de79

Observation ad36ac27-bea9-41e2-a4f1-d6857fd8fe50 · outbound

This paper cites NVIDIA Nsight Systems User Guide,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Systems User Guide,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.005213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.780615Z digest=sha256:88da70fb6e746f4b904f63d8f8799d27583794ee684d1ae150879c709e98a4b5

Observation 74e51260-d55d-4a1d-95e1-14039000aba4 · outbound

This paper cites Marconi: Prefix caching for the era of hybrid llms,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Marconi: Prefix caching for the era of hybrid llms,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.998021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.783043Z digest=sha256:992f99ba42ca7db41f00c4a4fc00f28f5952649b8012ec30ebb8e5a1650bd076

Observation 52486329-8c91-4d4b-84bd-4b94310c1a63 · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.990529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.788834Z digest=sha256:2752b66884b4b529e2d3c2ab2e441a9c55aa781366e01b87cb1fe77fbf3ba251

Observation 145fe7ed-7026-4bc8-a77f-92bca7747bd8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Qwen3.5: Towards native multimodal agents,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.982799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.791474Z digest=sha256:c08f47d2b27b904301a28a71a673449bebc8d5863baf05421d8b6d2e20f77c30

Observation ed74e44a-f1bc-4d3d-a066-6b49d9fe8740 · outbound

This paper cites Get to the point: Summarization with pointer-generator networks,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Get to the point: Summarization with pointer-generator networks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.974060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.793960Z digest=sha256:bf9e5d0e3b79ebf3c77463c720fef96224dcf74e34c15ff60a97ce311b24318f

Observation adb0e956-5cf9-4cc3-b540-22e1d91385f2 · outbound

This paper cites GitHub - sgl-project/sglang at release/v0.5.12 — github.com,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models GitHub - sgl-project/sglang at release/v0.5.12 — github.com,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.965305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.796907Z digest=sha256:f79776252267e9a3d57c6cecacb005966aeb3a90c7d5b6c07f25114457b40b3b

Observation 9eca6b96-d966-434e-b9cc-5227e861a25c · outbound

This paper cites anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.956695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.799558Z digest=sha256:8726c2cabee6eaf944309ecc874b6e4ce7928ec889a571287920818f66c6f629

Observation 27902834-7cb9-4369-9541-b26c00a13516 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llumnix: Dynamic scheduling for large language model serving,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.948954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.801941Z digest=sha256:c4416ec712f1225561bdf84239c1c85af11e9b35485bf05e99a17ec42a668100

Observation 2a5a3837-1995-4e42-b918-6191772cb2a2 · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.804544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.804544Z digest=sha256:28ca4d25ec4e56a6305d757362f1dc1199aa06fe1932f2111b6dbdbc51102221

Observation 74088bc6-074f-4dba-a5bf-651ef49861b7 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Triton: an intermediate language and compiler for tiled neural network computations,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.807486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.807486Z digest=sha256:b5edb0e3d6cc617eb621b351ff3ef5a5623dece995854ce505df81cab9decbf7

Observation 9f23af17-10ee-4ed3-ae8d-8f8dffa7159b · outbound

This paper cites Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.941600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.810674Z digest=sha256:1cf4b65156f1517b1517c18c09fd91b00c35dbd09ada847e20471f70a6f83455

Observation 9f2b2780-522c-4994-a2a1-e88fda4ec9ff · outbound

This paper cites Openhands: An open platform for ai software developers as generalist agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Openhands: An open platform for ai software developers as generalist agents,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.933584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.813367Z digest=sha256:542b3976738187ab08522ee9f73ec57da3f5fc4bb7f41a0e15a05a44d9fcbfd1

Observation 34f30b30-2b7c-4c4f-97c4-f776e3459680 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Roofline: an insightful visual performance model for multicore architectures,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.817103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.817103Z digest=sha256:23da4c267f3f09e8b5a302c20cabe80d11b899c82e3c7922f85ed9fd8daa9508

Observation 22984f73-11fc-4377-bf95-5170c7303300 · outbound

This paper cites Stree: Speculative tree decoding for hybrid state space models,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Stree: Speculative tree decoding for hybrid state space models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.924591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.819686Z digest=sha256:8e2cf9ebc45e270fdcc3e9ee3eb733ce390a8e750ca714586d15b1a50873eb6e

Observation defa3908-f307-47a4-8665-d722fd90f93b · outbound

This paper cites Strata: Hierarchical context caching for long context language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Strata: Hierarchical context caching for long context language model serving,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.915181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.822338Z digest=sha256:aa1fab8784b13af9b30d53e2126a215c76e9ddeae082d82e97d59002a5de8f6b

Observation da079084-1e09-4c0a-b498-92afe391c455 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Gated delta networks: Improving mamba2 with delta rule,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.906315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.825091Z digest=sha256:cc9bfa094de2978c1e43bda3225a1cb391d73b3dff8c22934120e203388ed0bc

Observation 9febaefa-f8f9-4399-89a0-a73dc2045386 · outbound

This paper cites Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.896471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.827799Z digest=sha256:1336d510ad7dc1c47cdd9fd437da6ddb1d11ec1fe24276512c0d6f470432b0da

Observation 37f2872d-6750-46d3-9319-69a66e55f7c5 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Flashinfer: Efficient and customizable attention engine for llm inference serving,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.887251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.830769Z digest=sha256:92dae5aa92063858958d783a09e552ba505a470d28e754e1d320c8de61aaf8ca

Observation 880a4d15-aa6e-4a15-9f48-a3ee80593255 · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.836278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.836278Z digest=sha256:bfcd1c86c862fc9e048e25a647d755b8e2434a3945459317e82749c5255bfa82

Observation 2dd1f0d1-d306-4d0a-a222-e9585658d688 · outbound

This paper cites Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,

Reference 51

Resolution
metadata mismatch
raw_fallback, observed 2026-08-04T23:34:56.952275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.838809Z digest=sha256:194db34a3793230e28df97fca566275f96e362406a67ae91c172af8dc304e3fb

Observation fa610865-6a1f-44f2-800d-443f12480b41 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.841217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.841217Z digest=sha256:686612c335e0db6659cc020805eca67390d736261c64b16c14aa37f306955254

Observation 4851dad1-8f2d-455d-be6b-1ab59ce296f0 · outbound

This paper cites Sglang: Efficient execution of structured language model programs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sglang: Efficient execution of structured language model programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.873171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.843854Z digest=sha256:2cf8544e0267759f685a000a5c26fa88b42c848d45237ccee52f7571e99b7c4e

Observation b65ea5f1-7791-4bf5-bc26-0800be44c666 · outbound

This paper cites NanoFlow: Towards optimal large language model serving throughput,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NanoFlow: Towards optimal large language model serving throughput,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.862730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:34:56.846403Z digest=sha256:7dd10dcce7637e5582e5152c5b125c42a0cc3530cea121c2c59a53d4b507353c

Observation 74e19205-6bed-40de-8329-6fe6d609e5ff · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.711032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.711032Z digest=sha256:47770a9e04fa16fd117c2b159d3ec8ef391da7cd0ba049fd7a2bc7564de5508b

Pith citing papers

No inbound Pith citation observations are available.