Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:02:39.419910Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 4 inbound Pith citation observations for arXiv:2506.03700.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:02:39.419910Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:16:08.552840Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T08:44:27.295523Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4bbd273e-ae5b-44cd-b242-8d013d4217c0 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e876ddf6-366d-4099-a62e-a32c355c1ba1 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0354e5-809f-4963-96b7-b4e489d0721f · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Hydra: Sequentially-dependent draft heads for medusa decoding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc87aa4a-3646-4859-b130-69d357c602d1 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Anthropic: Introducing claude 3.5 sonnet, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19c0759c-be96-4050-8989-a97eb1fd32cd · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Program Synthesis with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d12a47-2d5d-44d4-8aea-3aee6179fd5d · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8914505-6fde-4fe7-8f1d-0292bec2367d · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad9f962-f222-425e-8389-f2f82e441797 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d50787-59f9-4d9b-9fde-cbfb998908a9 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism D., Chen, D., and Dao, T
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6f76329-635d-4bdd-8141-deb9541e6a9e · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating Large Language Model Decoding with Speculative Sampling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517a3181-d6a9-43b6-b6a7-e03263455086 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism WAPITI: A Watermark for Finetuned Open-Source LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2995c851-694e-414b-8043-b2cf4dc76ec0 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Evaluating Large Language Models Trained on Code
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06865bfe-f5cb-4986-bc51-0c98ec6ac986 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Verifiers to Solve Math Word Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ab56a2-a86e-48d8-aeee-ffb48ab8ecca · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Flashattention-2: Faster attention with better parallelism and work partitioning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a98939-fdc4-4245-b86f-46aa7c54c832 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e49189-a87c-4c21-9a4e-6006e3a73a7f · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8387e030-7012-406e-b94c-0a712bcafefc · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38732595-4509-47ae-922b-34fc6e2a3a39 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism QLoRA : Efficient finetuning of quantized LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b6071e4-831a-43f8-8200-da6fc141262c · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab443d8-4c37-4ed2-881b-a094f71fe195 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Glide with a cape: A low-hassle method to accelerate speculative decoding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15e11d1e-e8f8-4514-815f-f1ae43ebe166 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The Llama 3 Herd of Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86cd2b5-5a9a-440d-be4c-034879542cfd · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Depth-adaptive transformer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a80e998c-4443-4ea8-adbe-f0323b900831 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L ayer S kip: Enabling early exit inference and self-speculative decoding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 201dc01e-7564-4e6e-a3fa-9b6a8a11b585 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Break the sequential dependency of LLM inference using lookahead decoding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb14de1f-0ab4-4afb-8f6a-818d935ffdca · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94f56415-7df1-433c-b465-91a8350e4ce4 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bed347e-7dc9-447c-a58b-b38a6354fe65 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism REST : Retrieval-based speculative decoding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a12cfdc8-b27b-4192-a2d0-6e48d50e6eb5 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Training Compute-Optimal Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daff1876-b4f7-4876-b4fd-120090b0c010 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism The curious case of neural text degeneration
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b54c427-bdf2-4722-a5b0-6f787aff1049 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SPEED: Speculative Pipelined Execution for Efficient Decoding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8031e3f1-13b2-4eed-8dad-9b94ecfa7908 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92faf0a8-b416-4acc-8d08-63636b1f2689 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Efficient Test-Time Scaling via Self-Calibration
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b37d04a-8585-4b7f-9845-86a22812cabb · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Multi-scale dense networks for resource efficient image classification
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6ba6a95-54c5-4eb3-b122-6a9b5d7193e7 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757b7c6c-b126-4ce0-84d0-6d6d96164eeb · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 411e6a88-6391-4fd9-b649-592e0f0e39da · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixtral of Experts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef25ad9-1f64-4dba-a51a-44173e48128e · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sigsoftmax: Reanalysis of the softmax bottleneck
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b63bef3-2853-4021-818a-d4ccc8d5b4e0 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling Laws for Neural Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 471210c7-f1bc-4815-8654-554b09d948a9 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2beb81a9-c3c8-4764-a4ce-583a06234c48 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism W., Gholami, A., and Keutzer, K
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7f8bf5-bb5c-4c31-981c-8a06f41d05cf · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef55932-148c-41f3-9398-670a7691f026 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Adam: A Method for Stochastic Optimization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9866707e-911f-403d-8b81-2558603b726e · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Fast inference from transformers via speculative decoding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94acec46-634f-44f3-b2e3-ff75c53c7aaa · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE : Speculative sampling requires rethinking feature uncertainty
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c182b05-b6c8-49b4-9489-3bab0d4c6b74 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism EAGLE -2: Faster inference of language models with dynamic draft trees
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a3730a1-1a9a-448b-b099-e8041ad0a0d7 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e6d8d3-1059-4a40-b3e0-6c887e8ac97c · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Speculative decoding via early-exiting for faster LLM inference with T hompson sampling control mechanism
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 392f706e-b5d3-4f78-8388-c4441d6cd354 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Online Speculative Decoding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d65e517c-e34b-44ef-8000-15ee09bcc21e · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b92ac1-0853-49ca-b29a-7b46567ace67 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524c374e-19b6-41cb-94cf-db6ad6372692 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbeac526-b188-40c9-92ca-a53d9fd3eff2 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism B., and Lapata, M
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6dbb81cc-5103-41ad-b311-0f9705e9b60f · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Introducing OpenAI o1: Learning to reason with large language models, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5b52e0b-3b03-412f-8a18-be2363e6336d · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Suri: Multi-constraint Instruction Following for Long-form Text Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 137b5722-76e6-45f9-93f9-7dd36980ffda · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d2b5b6-ba6f-4ae0-bd91-d6335a15b7d7 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Zero: Memory optimizations toward training trillion parameter models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f1598f-a45e-45c9-8378-1ca46fde6337 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5b092e6-bafb-4015-b99a-ff921311b22b · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Confident adaptive language modeling
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a47a6e7-9004-4fb9-af90-21d5e8a57ec0 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b9a603-75a1-4c5a-9dc0-d52cb414d087 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Blockwise parallel decoding for deep autoregressive models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90a9b21-851c-4a9f-957b-736288df6d89 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Branchynet: Fast inference via early exiting from deep neural networks
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e438b8-eb7b-48b9-9186-9aea743e3c0b · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Accelerating LLaMA Inference by Enabling Intermediate Layer Decoding via Instruction Tuning with LITE
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2d3b77-700d-4849-9bce-2a0b5b5c1298 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a266fe9-7a37-46dd-ac7d-0d2d051c1307 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SWIFT : On-the-fly self-speculative decoding for LLM inference acceleration
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c447bb5-cafd-47be-b11f-5fc3ac132133 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Sheared LLaMA : Accelerating language model pre-training via structured pruning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b1fd645-bab9-4770-a385-c6b917c22b62 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Predictive pipelined decoding: A compute-latency trade-off for exact LLM decoding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21a3dae6-973a-4857-968f-03fbbd74d4bf · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99ec4d90-e5a2-4bc7-9b00-edef53252766 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism C o S afe: Evaluating large language model safety in multi-turn dialogue coreference
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ab76d82-5c3e-4c69-b777-6e309c525fb3 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b43b1424-d65b-4f94-b3e4-c5ea393af016 · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism Draft & verify: Lossless large language model acceleration via self-speculative decoding
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98a08fef-5fd2-4b5e-a4e8-127f583d7fcc · outbound
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism SafetyBench : Evaluating the safety of large language models with multiple choice questions
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6daaf60-f3ba-4e44-b348-d33c26c2b93a · inbound
HiSpec: Hierarchical Speculative Decoding for LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475e5a31-0525-44da-9544-e3508eba031d · inbound
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 642cdc18-96fd-4579-97ce-916b42d78cd2 · inbound
Depth Exploration for LLM Decoding AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca429a9b-ae99-4be2-b3ba-6fe0e5517d08 · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.