Pith. sign in

Paper Citation Record · LEDGER

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 4 inbound Pith citation observations for arXiv:2506.03700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03700 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:02:39.419910Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:16:08.552840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:44:27.295523Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4bbd273e-ae5b-44cd-b242-8d013d4217c0 · outbound

This paper cites write newline.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.111163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.111163Z digest=sha256:ec4953b47d8b9fbf4ed3bb236b28a6e576a5fd6d5aaaee84a5483a934c2e8174

Observation e876ddf6-366d-4099-a62e-a32c355c1ba1 · outbound

This paper cites GPT-4 Technical Report.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.117325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.117325Z digest=sha256:2dac143b7c327e28ed87214405537be61860c337c8e34ca667870010f7fac83f

Observation fa0354e5-809f-4963-96b7-b4e489d0721f · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Hydra: Sequentially-dependent draft heads for medusa decoding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.542572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.122576Z digest=sha256:35301dee01780b869e6554e9b451894b0735eb83d53d1e0b1155ef34b576367c

Observation cc87aa4a-3646-4859-b130-69d357c602d1 · outbound

This paper cites Anthropic: Introducing claude 3.5 sonnet, 2024.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Anthropic: Introducing claude 3.5 sonnet, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.526458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.127205Z digest=sha256:469d6d736f28cbfb29396c28fd19c50c7c51583d324143c31946debbb9a2f2b2

Observation 19c0759c-be96-4050-8989-a97eb1fd32cd · outbound

This paper cites Program Synthesis with Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.131952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.131952Z digest=sha256:2e70c59afba358ddc02dfcf2ffdc68dc451e9bca507be125dd61de258169abc2

Observation 19d12a47-2d5d-44d4-8aea-3aee6179fd5d · outbound

This paper cites Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.136964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.136964Z digest=sha256:83a02bf2f896707e3a00ca0f51007556e7b6b3ff9e27d8b22aded4ba6fae237f

Observation b8914505-6fde-4fe7-8f1d-0292bec2367d · outbound

This paper cites LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.141519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.141519Z digest=sha256:534b02ba1c8cab58429b2e277004f2ac442f1ceaee57eac6d55ddda127e53f0f

Observation dad9f962-f222-425e-8389-f2f82e441797 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.146799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.146799Z digest=sha256:65938b443b05471a396b6e95aa1a77d709cb2b5bed05fda8b73c68aa37a77b29

Observation a6d50787-59f9-4d9b-9fde-cbfb998908a9 · outbound

This paper cites D., Chen, D., and Dao, T.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism D., Chen, D., and Dao, T

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.499957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.151371Z digest=sha256:a780a792c909b04dd2a2b315b8be5f0044d3077710b92332578c8003f665dc92

Observation d6f76329-635d-4bdd-8141-deb9541e6a9e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating Large Language Model Decoding with Speculative Sampling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.155905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.155905Z digest=sha256:b1f1a264364662a43be27c88b322803b6c066779fcaca1261ca77d83f3899ac1

Observation 517a3181-d6a9-43b6-b6a7-e03263455086 · outbound

This paper cites WAPITI: A Watermark for Finetuned Open-Source LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism WAPITI: A Watermark for Finetuned Open-Source LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.160378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.160378Z digest=sha256:7be05b33e0eef5b40a644f6130bef775bf58548fd3d85443f79165a80f0d5df6

Observation 2995c851-694e-414b-8043-b2cf4dc76ec0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Evaluating Large Language Models Trained on Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.164777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.164777Z digest=sha256:b02eb26ccad0324ac3db43be4274af9b2e308246c5d196ca4395f6ec674267a3

Observation 06865bfe-f5cb-4986-bc51-0c98ec6ac986 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.169280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.169280Z digest=sha256:ef97893d45985010c234e86086f5eb5fdaee9eb7c4e01e70480cd504615d8eaf

Observation b5ab56a2-a86e-48d8-aeee-ffb48ab8ecca · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.173568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.173568Z digest=sha256:05538948bd679b289e4b1e8edbefd8430699396ab28c2fbac424102e0ec53968

Observation c5a98939-fdc4-4245-b86f-46aa7c54c832 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.177936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.177936Z digest=sha256:50f61727cffcdb1a5d395a07121fd69b68a5f17e82ed485a72c6ac5f966b021f

Observation 18e49189-a87c-4c21-9a4e-6006e3a73a7f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.182650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.182650Z digest=sha256:e083ee2d29360f467acf35100bfff95f3ad605ad4e0603a966e024bcb0ad2a31

Observation 8387e030-7012-406e-b94c-0a712bcafefc · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.187286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.187286Z digest=sha256:4973f7694d630443a8a3bdfabbb38c72e263931edec608b9b49631e0f6ebf897

Observation 38732595-4509-47ae-922b-34fc6e2a3a39 · outbound

This paper cites QLoRA : Efficient finetuning of quantized LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism QLoRA : Efficient finetuning of quantized LLMs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.474491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.193356Z digest=sha256:e0651c713462642c59a91f17f3863cb30bceaec75a48ee6feb3d5c5901058339

Observation 7b6071e4-831a-43f8-8200-da6fc141262c · outbound

This paper cites Jump to Conclusions: Short-Cutting Transformers With Linear Transformations.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Jump to Conclusions: Short-Cutting Transformers With Linear Transformations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.197592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.197592Z digest=sha256:aef577408843231d3ca9085d979e0289394ab9e45caed5ed83811518258749e2

Observation 8ab443d8-4c37-4ed2-881b-a094f71fe195 · outbound

This paper cites Glide with a cape: A low-hassle method to accelerate speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Glide with a cape: A low-hassle method to accelerate speculative decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.459979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.202238Z digest=sha256:b01af4cddd27a28cb8ff491cdeb3b366269b2ddb816c97c79d8655e7f4b0c195

Observation 15e11d1e-e8f8-4514-815f-f1ae43ebe166 · outbound

This paper cites The Llama 3 Herd of Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.206389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.206389Z digest=sha256:2f1757559c9777105cea85406869b44e91fc45fa2a24d98e25b5650f06b7c5ae

Observation e86cd2b5-5a9a-440d-be4c-034879542cfd · outbound

This paper cites Depth-adaptive transformer.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Depth-adaptive transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.445245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.210467Z digest=sha256:418bc0611a2ba6ac568264f57d58eb86c6b0d7581cdeeb3b46cfd5bba62615cc

Observation a80e998c-4443-4ea8-adbe-f0323b900831 · outbound

This paper cites L ayer S kip: Enabling early exit inference and self-speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L ayer S kip: Enabling early exit inference and self-speculative decoding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.430525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.214544Z digest=sha256:e8a9b6fd9aad02ad6dc551b8e0b440f13109d5c9895c5371d63a86390d070af1

Observation 201dc01e-7564-4e6e-a3fa-9b6a8a11b585 · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Break the sequential dependency of LLM inference using lookahead decoding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.415153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.218593Z digest=sha256:84b259c4d44e84b0851b68df5cf1a185eae82e8cfc7036b53e7ca9a23f2261f2

Observation bb14de1f-0ab4-4afb-8f6a-818d935ffdca · outbound

This paper cites Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.399611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.222519Z digest=sha256:626bbace0627166681870834ac07f8a3748aa212543c16fee260e7e35e800404

Observation 94f56415-7df1-433c-b465-91a8350e4ce4 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.226807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.226807Z digest=sha256:f384341f77e9f2626cc4c20f1cd6dcc121622165c842a7a262ef230e20876dbe

Observation 5bed347e-7dc9-447c-a58b-b38a6354fe65 · outbound

This paper cites REST : Retrieval-based speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism REST : Retrieval-based speculative decoding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.383163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.231129Z digest=sha256:0c684f8ea926a41cb873fa7457f3ee6c5bb60550dc9b86e15931d79707e6d6da

Observation a12cfdc8-b27b-4192-a2d0-6e48d50e6eb5 · outbound

This paper cites Training Compute-Optimal Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Compute-Optimal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.235113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.235113Z digest=sha256:5bd1313737f3863e00f9e8b2e479941c22dad18b5e1d4219820610a95cf73ba2

Observation daff1876-b4f7-4876-b4fd-120090b0c010 · outbound

This paper cites The curious case of neural text degeneration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The curious case of neural text degeneration

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.366794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.239313Z digest=sha256:38788592a5d3684bdd8227963a2421b691636620aa79f9d7aed4a2dd72ef718b

Observation 5b54c427-bdf2-4722-a5b0-6f787aff1049 · outbound

This paper cites SPEED: Speculative Pipelined Execution for Efficient Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SPEED: Speculative Pipelined Execution for Efficient Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.243484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.243484Z digest=sha256:14211004e0c2fea3f235e90186ee9256df9a5e115c1a3c33b9f2c5cf0554fedf

Observation 8031e3f1-13b2-4eed-8dad-9b94ecfa7908 · outbound

This paper cites E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.351586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.247841Z digest=sha256:8dfad3c9614cc836249022140d2faa896c27dc9156d3076e436df089ee28f7ee

Observation 92faf0a8-b416-4acc-8d08-63636b1f2689 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Efficient Test-Time Scaling via Self-Calibration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.251956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.251956Z digest=sha256:76c96e3306c01426fef1db1a005abafa2517c0419ae88057f1f5411a85a63181

Observation 0b37d04a-8585-4b7f-9845-86a22812cabb · outbound

This paper cites Multi-scale dense networks for resource efficient image classification.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Multi-scale dense networks for resource efficient image classification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.336787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.256248Z digest=sha256:c0c663f3af0e23974761a6e9f7108c4aacc9e5d44f1e541e5ce215bd2c9d752f

Observation c6ba6a95-54c5-4eb3-b122-6a9b5d7193e7 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.260580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.260580Z digest=sha256:194c4dc5d7bafa5c641f719fd782eb31418b0c39645ab103f197cf1ca63ebd4e

Observation 757b7c6c-b126-4ce0-84d0-6d6d96164eeb · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.321949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.264904Z digest=sha256:74d5421c667af74a412c88723b0ec47aaabb876be213b0c3373de1bcb87055b3

Observation 411e6a88-6391-4fd9-b649-592e0f0e39da · outbound

This paper cites Mixtral of Experts.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixtral of Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.268868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.268868Z digest=sha256:689a62c8dd08be13e0e2ccef275ef3d93d197d74b3a98bb47e706aacf9d7bc29

Observation 8ef25ad9-1f64-4dba-a51a-44173e48128e · outbound

This paper cites Sigsoftmax: Reanalysis of the softmax bottleneck.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sigsoftmax: Reanalysis of the softmax bottleneck

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.307036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.273111Z digest=sha256:ac9c37883705d546b50737779a939e5263b4a4d0f03516b02f287438b4c321a4

Observation 1b63bef3-2853-4021-818a-d4ccc8d5b4e0 · outbound

This paper cites Scaling Laws for Neural Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling Laws for Neural Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.277280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.277280Z digest=sha256:6fcab1feb433a86910782d8f782be545b45d4218a5837661fc68a39bc683a344

Observation 471210c7-f1bc-4815-8654-554b09d948a9 · outbound

This paper cites A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.281273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.281273Z digest=sha256:acc1e316eed67836935c69ea3146c3cf2b96ed4feef59c8c6d164559cdc99e7c

Observation 2beb81a9-c3c8-4764-a4ce-583a06234c48 · outbound

This paper cites W., Gholami, A., and Keutzer, K.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism W., Gholami, A., and Keutzer, K

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.285634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.285634Z digest=sha256:4df83766c3f04906b02d8b2e8edb0a46897e6664d2dbadc8914467626d8795d6

Observation bb7f8bf5-bb5c-4c31-981c-8a06f41d05cf · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.289715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.289715Z digest=sha256:5618b816bee130303c41154f123a5be937669c6cf74e3b8563f1c493241c6d0b

Observation 3ef55932-148c-41f3-9398-670a7691f026 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Adam: A Method for Stochastic Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.294461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.294461Z digest=sha256:38a98568b23c9f94d52f54fce98cef44204e033ae29d18412c6468efb77b3a65

Observation 9866707e-911f-403d-8b81-2558603b726e · outbound

This paper cites Fast inference from transformers via speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast inference from transformers via speculative decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.298736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.298736Z digest=sha256:1d02c7494f17932ac6e10c01e025316b268cf8170ac55c98ef248e67014ae876

Observation 94acec46-634f-44f3-b2e3-ff75c53c7aaa · outbound

This paper cites EAGLE : Speculative sampling requires rethinking feature uncertainty.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE : Speculative sampling requires rethinking feature uncertainty

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.268093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.302878Z digest=sha256:ec4831b1f22e7d4417071ff721cd688dcefe795f948b42e417bf4391309ce0c1

Observation 8c182b05-b6c8-49b4-9489-3bab0d4c6b74 · outbound

This paper cites EAGLE -2: Faster inference of language models with dynamic draft trees.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE -2: Faster inference of language models with dynamic draft trees

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.252446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.306771Z digest=sha256:7fb93604217f44ef9679d5b29324b5907bf0e96daee812f3c6ec898fd99213e4

Observation 4a3730a1-1a9a-448b-b099-e8041ad0a0d7 · outbound

This paper cites Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.310757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.310757Z digest=sha256:c4b5f6631a3d6af31612532b03d754903b15b774930f612ee6d0d6d682a3a2a7

Observation c3e6d8d3-1059-4a40-b3e0-6c887e8ac97c · outbound

This paper cites Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.236346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.315020Z digest=sha256:e418f04938c1c10b725d46dcfca1eba71fa16302bc5a118baca4cdaa3ed12c21

Observation 392f706e-b5d3-4f78-8388-c4441d6cd354 · outbound

This paper cites Online Speculative Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Online Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.319649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.319649Z digest=sha256:d590ec491dad24225e129aac5339c71b06be60b4a655f03da7127adc9cde7dbb

Observation d65e517c-e34b-44ef-8000-15ee09bcc21e · outbound

This paper cites SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.324065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.324065Z digest=sha256:ea04bc6945c9a9eed46e25ac2bde09246bb9b73ff12d4e1f625bb51b2a6f3953

Observation 13b92ac1-0853-49ca-b29a-7b46567ace67 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.328762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.328762Z digest=sha256:4760a6c889e6d11258134bfedb22959c5bd3d82509927f4a208d5fad80476d9d

Observation 524c374e-19b6-41cb-94cf-db6ad6372692 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.332990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.332990Z digest=sha256:8a4d442e1526954a6f46e5cadaa8d989c0e04590746cef1ac660edc1864ecad6

Observation fbeac526-b188-40c9-92ca-a53d9fd3eff2 · outbound

This paper cites B., and Lapata, M.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism B., and Lapata, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.220011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.337310Z digest=sha256:7e14ca9a887a1ffdabae42c3c5d12c38abbff56e8e68f0de3084d26079e0f5e5

Observation 6dbb81cc-5103-41ad-b311-0f9705e9b60f · outbound

This paper cites Introducing OpenAI o1: Learning to reason with large language models, 2024.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Introducing OpenAI o1: Learning to reason with large language models, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.205186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.341535Z digest=sha256:0635e36da4bda6d2dab84703942f3fdc01889a9e69c36b5d586aa9ef42633f8c

Observation b5b52e0b-3b03-412f-8a18-be2363e6336d · outbound

This paper cites Suri: Multi-constraint Instruction Following for Long-form Text Generation.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Suri: Multi-constraint Instruction Following for Long-form Text Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.345559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.345559Z digest=sha256:917c4681f8184c3630903de82ce3b14afd21367ba60f16ce4324378c1aecf6ea

Observation 137b5722-76e6-45f9-93f9-7dd36980ffda · outbound

This paper cites Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.349604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.349604Z digest=sha256:cdb758b1fe556b1b74adf90d5d24753552c661b242050643dc5dc6178a94a8fc

Observation c7d2b5b6-ba6f-4ae0-bd91-d6335a15b7d7 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Zero: Memory optimizations toward training trillion parameter models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.354357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.354357Z digest=sha256:d576b409f20401eddc51b9224894041ec3c93db9ed214d8f9b9e639ea8cfc904

Observation 00f1598f-a45e-45c9-8378-1ca46fde6337 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.358502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.358502Z digest=sha256:4aeb48db7908b57ff8c842be80b300aec958b8f4c303673baffdd9309c71615f

Observation c5b092e6-bafb-4015-b99a-ff921311b22b · outbound

This paper cites Confident adaptive language modeling.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Confident adaptive language modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.363381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.363381Z digest=sha256:49ca01dd6f699eca4f04773b41dd52531cbf742357fe6211cf79e4bc9070a547

Observation 1a47a6e7-9004-4fb9-af90-21d5e8a57ec0 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.367516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.367516Z digest=sha256:23b6462e79a0921bb21de2ace2950a0f934886efc14e67f41932460df8571943

Observation 80b9a603-75a1-4c5a-9dc0-d52cb414d087 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Blockwise parallel decoding for deep autoregressive models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.371960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.371960Z digest=sha256:a63e27a1d889ff90972929529d7df2126d01c7cba7d554bab58f9bad22818387

Observation f90a9b21-851c-4a9f-957b-736288df6d89 · outbound

This paper cites Branchynet: Fast inference via early exiting from deep neural networks.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Branchynet: Fast inference via early exiting from deep neural networks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.376015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.376015Z digest=sha256:838fe0ed8ba6aede5964de4a51a94fecd73919452f6cf893e42339bd0b8b5211

Observation 91e438b8-eb7b-48b9-9186-9aea743e3c0b · outbound

This paper cites Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.380095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.380095Z digest=sha256:6a0a65a4ba873189dd80b43f32a61c8afa30d9dd8e4ef41fc6787bb12e2eef3b

Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:02:39.384667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:02:39.384667Z digest=sha256:d529b102ae1c8ce77bcc920daddba12dcc5fffbc35718af070c3a427832c1f7f

Observation 6a266fe9-7a37-46dd-ac7d-0d2d051c1307 · outbound

This paper cites SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.146961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.389151Z digest=sha256:e9230282a709bc899af089c0eca23871cafca24632ac83fdba753d46eadd606b

Observation 3c447bb5-cafd-47be-b11f-5fc3ac132133 · outbound

This paper cites Sheared LLaMA : Accelerating language model pre-training via structured pruning.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sheared LLaMA : Accelerating language model pre-training via structured pruning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.131589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.393390Z digest=sha256:d0023ce5d6fce329304e0bde7c441206996c8135900553f4c65afa81251f9499

Observation 1b1fd645-bab9-4770-a385-c6b917c22b62 · outbound

This paper cites Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.114487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.397839Z digest=sha256:005cbd7c75330ae6737319b6d1854e10b0b86e3da1e33c0a8bb91e5e2fcc1f5f

Observation 21a3dae6-973a-4857-968f-03fbbd74d4bf · outbound

This paper cites an unresolved cited work.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:02:40.097652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.402276Z digest=sha256:bb3f7740f18068179e03dfd3aed0d5fafc660b70c553c0b74b22b39a547c8574

Observation 99ec4d90-e5a2-4bc7-9b00-edef53252766 · outbound

This paper cites C o S afe: Evaluating large language model safety in multi-turn dialogue coreference.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism C o S afe: Evaluating large language model safety in multi-turn dialogue coreference

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.082567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.406543Z digest=sha256:eca5bb5fa4b55052fe681a11c1322798e8d25f3f09ce4f5b86fa46fd4712421a

Observation 1ab76d82-5c3e-4c69-b777-6e309c525fb3 · outbound

This paper cites L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.066657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.410964Z digest=sha256:dc60cfe5355e2f9ee6a02b260becb5ead2e46507aaa15f89b2d389aebad97794

Observation b43b1424-d65b-4f94-b3e4-c5ea393af016 · outbound

This paper cites Draft & verify: Lossless large language model acceleration via self-speculative decoding.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Draft & verify: Lossless large language model acceleration via self-speculative decoding

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.051401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.415285Z digest=sha256:d2d9c78dff50e1968f047193b61af655c1fb5489ea3805a43661c44fdfe68f0c

Observation 98a08fef-5fd2-4b5e-a4e8-127f583d7fcc · outbound

This paper cites SafetyBench : Evaluating the safety of large language models with multiple choice questions.

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SafetyBench : Evaluating the safety of large language models with multiple choice questions

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:02:40.035083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:02:39.419910Z digest=sha256:3fcee50151cb0c72a1e1fbe48622d09251186bd47d6956dbc4189b434f656a2c

Pith citing papers

Observation f6daaf60-f3ba-4e44-b348-d33c26c2b93a · inbound

HiSpec: Hierarchical Speculative Decoding for LLMs cites this paper.

HiSpec: Hierarchical Speculative Decoding for LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:30.182708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:30.182708Z digest=sha256:19a29a0694ee12740b672df31cf6e2ef5c5d7fb3e3e83f155c774f60813bcabb

Observation 475e5a31-0525-44da-9544-e3508eba031d · inbound

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration cites this paper.

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:58.697801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:22:05.049712Z digest=sha256:167d96a52cec120eb932c410cb23f4bc18ff5e46c348d1098d22feda9263580c

Observation 642cdc18-96fd-4579-97ce-916b42d78cd2 · inbound

Depth Exploration for LLM Decoding cites this paper.

Depth Exploration for LLM Decoding AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:44:27.297075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T08:43:47.472682Z digest=sha256:4a4b576a9dca3fc1133759125e1dc5b2271678d5b1a6e4a7ec39f523852a56ef

Observation ca429a9b-ae99-4be2-b3ba-6fe0e5517d08 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.552840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.552840Z digest=sha256:7dd88c4fceb61763434a28d34df7b1453b569735962b96ffb9f8201d7d8a1737