Pith. sign in

Paper Citation Record · LEDGER

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2509.04474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04474 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:48:23.144764Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0f3bf10-0159-4c65-a2a0-109be6e5d20d · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.159265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.159265Z digest=sha256:10603f7727b5dc96a8298b8faa9d3e7a6fd0e69221bf1fe241d61bad1f45c7cc

Observation fda6280c-cd33-442d-8336-8d4c82be5159 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Training Large Language Models to Reason in a Continuous Latent Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.516547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.516547Z digest=sha256:4e22754dc424d3c934b3ae6875bc8ddb0b38c3f4fbe07f3abba4a872c7d22330

Observation 97d15c73-75eb-468f-b5a4-b949fa073faf · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.574388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.574388Z digest=sha256:735733e7eeda186b2d5de9a087d2d8c5550379dea6e43cbd3a0af0b16124e789

Observation e8487933-93ff-4006-9f1e-1d273b3ac010 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Competition-Level Code Generation with AlphaCode

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.656281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.656281Z digest=sha256:df80c89dc35372263c2557e52fe7e86e0e40857cb3ca60f78ea0b04e746235bf

Observation e2392494-a000-412a-a0a1-822dcb38bf9e · outbound

This paper cites Let's Verify Step by Step.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.885995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.885995Z digest=sha256:89ecb3fc2012ce2ee303bb0be635c7d960ca628e7c7c7748516a5d19596619a2

Observation b75ec70c-a409-46f9-af64-e1c149f1ad54 · outbound

This paper cites TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.999552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.999552Z digest=sha256:5edd25eaba6dbe0f66ff3c5ad25c92ccdc0f362b25db52ea7f9913bc7600b583

Observation 5498e677-3537-4449-a6be-a819a8f98388 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Reasoning Models Can Be Effective Without Thinking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.081867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.081867Z digest=sha256:fd1c0a5358007de0ffe169d7c2bf57bb4d2cbca2ba163f01ad4c5029c168f492

Observation 652600af-233b-4e66-b3b0-db8b68ae23de · outbound

This paper cites Suffixdecoding: Extreme speculative decoding for emerging ai applications.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Suffixdecoding: Extreme speculative decoding for emerging ai applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.195104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.195104Z digest=sha256:10dc26d35affeefd7c7ca71d6f88ef57d73248541eeed67e7c6056df4d27c6e4

Observation e740dd57-f9fe-439c-b562-be754c3a9e5e · outbound

This paper cites OpenAI o1 System Card.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.337886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.337886Z digest=sha256:5f9beddd86c305596f23ba784495e8872906791d2af6bd00d668ba3bd6a0fe0d

Observation 5774eab7-57c0-4983-a394-583f0c9d3f03 · outbound

This paper cites SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.435299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.435299Z digest=sha256:3347321c9bbf85a4a3394ebd8d6d2a4b2d5d78d905542816c0a9ab26b9d05a96

Observation d1a4389e-4c0a-4ff9-a3e4-15cf770075bc · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.547651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.547651Z digest=sha256:70d7ae3a77af92ce33d01bff41f5963248d1a78269eb70f269af855ea674e328

Observation fc50e29f-9b0c-4dff-9163-54b06bb2b2c0 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.659104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.659104Z digest=sha256:f3fbf723a3800a0372b332e26470dda084f8be02b133f21c57c37ab87730bd87

Observation 248c6c22-6920-42e5-9b6a-c6d3a7e41172 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.756427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.756427Z digest=sha256:de83304deb0d062a37bf5126e1982ca865d10486aab03e370ac39f4882d9a721

Observation 66927755-e89d-481b-97db-e4e91bf8aeeb · outbound

This paper cites R1-compress: Long chain-of-thought compression via chunk compres- sion and search.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling R1-compress: Long chain-of-thought compression via chunk compres- sion and search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.822091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.822091Z digest=sha256:44d1aa53e89726e2bdfd5d893bdc187e158ce58847c115d0c8c2cc31bac0fec2

Observation 178cae99-eb4f-45a1-9c87-f08a431ad679 · outbound

This paper cites Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:22.954670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:22.954670Z digest=sha256:255da555262f2c9b3ba36865277f06fcfc5c0347f06b9853e7a7d6df5bff8da8

Observation d24ca8c5-ad1c-4519-b0ba-d3b3ceea96cf · outbound

This paper cites Qwen3 Technical Report.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:23.073663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:23.073663Z digest=sha256:770eaf1596428dfe758f4bf7c82aae329205c9e2626e878e2db4449d6de9f7e1

Observation 5ff7f1ea-4082-41c1-90fb-2770e16904bf · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:23.144764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:23.144764Z digest=sha256:6f324f3fb8ec7b298a98bf4f79d874a68f6cc28b90b0a8bf280ee8643cd91c23

Observation 393a2bd1-c2b2-4eab-a223-8903dcaf506e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.341874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.341874Z digest=sha256:0fbdadad09d5f3f210ef30778d21e038a702fade5392592cee84dbd2e4717a3a

Observation 591a950f-0fcf-4023-b754-e1aa1b44acc3 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.764917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.764917Z digest=sha256:f82da3a8c6bbb15a7ed39f8f2073258a703f6275034bc9097e0a919b6bdc2a09

Observation eb14ed75-474b-4cd7-87a3-cce4ce4bebb3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.248460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.248460Z digest=sha256:e426ba08691b345d25b29a1895efe4bb4de1952c8d953f09002599dc7e9960d5

Observation ef734227-f47c-4a37-9275-0c36f298f419 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.101641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.101641Z digest=sha256:5fd516c81490d660d2c66bf507d01c0156f425ac9bcb3e5482beb7fd3c4cd5d8

Observation c386ed0c-2dca-44de-98e1-0ab2d484f2ae · outbound

This paper cites Break the Sequential Dependency of LLM Inference Using Lookahead Decoding.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.420932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.420932Z digest=sha256:bc55a8729676b50e557861e0555fdcb0d2d1d0e534f480f06291902c09e733ed

Pith citing papers

No inbound Pith citation observations are available.