Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:23.901164Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2506.22396.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:23.901164Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-14T22:59:07.065193Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-14T22:59:34.279645Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cbdea40f-a78e-4bf2-a233-45def0e3706a · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This ": processed all layers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 55d60984-3f8a-4332-ad5b-83b8aa8f4132 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization this ": KV diff 1.00 < 0.30 -> Write
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ceffd707-cb7d-43bb-a4ab-7af4522662a4 · outbound
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f923ebb3-9e92-4cd0-9245-b8626840ec47 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc147b7-69e6-4816-baf2-79182af41f32 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d563ab34-0ddb-4def-8e66-9e57cdad4304 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Adaptive Computation Time for Recurrent Neural Networks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10675910-233b-4717-b9d8-d1a0a413bb93 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e1305a7-c7ab-4ee0-80cc-4a5aeb0c3cb4 · outbound
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9830d11b-602f-42ed-966e-906875262a7b · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Transactions of the Association for Computational Linguistics, 8:842–866
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3005d7e8-64b6-4768-9db1-f58c93b9e6b5 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization and ": entropy 0.23 -> 2 - bit quant Token
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5de69b14-c194-45c1-8a33-36faf45a01f1 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 70b860aa-9582-4900-a298-c52495bfbb84 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be1bc39a-c227-454f-9c9b-023fc1d7c74b · outbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 340a6210-5183-4572-bc9e-972aa58128c8 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8e967772-e3fd-432c-971b-ea39e29bc34a · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b44abbd4-911d-4b2d-8854-64d7dd03a1fb · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization skipped after Layer 2
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac9c5c76-258e-483a-9422-72051c13e320 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 50b1b291-d620-430c-afe5-a794443dc6fb · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ff59b215-10fa-4728-a5df-f16c00033a54 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 320abb27-1600-4517-9b8f-8980131a5ebe · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This conservative threshold ensures fusion only when representational collapse is semantically safe
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 13489094-6c33-4013-9521-78b40884a72b · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3a64abf-0ac8-46a2-bd14-9a857917d78e · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization We choose the pair (τlow, τhigh) = (0.3, 0.6) that achieves a strong Pareto frontier: ∼ 8.6% additional FLOP reduction with less than 0.1 perplexity change
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dff6614e-5875-45ae-b55d-43a4879dbd19 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9b1963e1-8a94-4e7c-b5d8-77fe94faa165 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5780c87f-a0be-4ef7-aac8-637ceecbbc95 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Pareto Surface Analysis
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2632d43f-6fdb-40dc-a5a3-e86ccce7e131 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a48624b-df50-437f-a210-2fa41d532593 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 109bdff8-33a9-4348-a66a-093ce5cc00fb · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization We hope this work encourages the community to adopt tools like CodeCarbon not as post-hoc profil- ers, but as first-class citizens in the deployment pipeline
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 21bf7a74-74a1-40a4-88b5-0dc614ac3f85 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization I.2 holds, halt the token unconditionally
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f17e496-1fa0-4250-a091-099349017816 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization I.3 holds for any u, merge (t, u)
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a014c21-1968-4921-a1bf-73f7835c4458 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This priority is grounded in halting the provision of computational savings without representational loss, while fusion entails approximation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a247a1bb-80cf-4184-a20b-6a4175ef868f · outbound
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53f694ef-434c-44fb-8cba-210f44d7265b · outbound
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 049f2228-6a3f-4ccf-86d7-8905790b5e24 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Shaojie Bai, J Zico Kolter, and Vladlen Koltun
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 76f29783-9953-4308-a35a-b0f1be593238 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Fast Inference from Transformers via Speculative Decoding
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f82ae740-22a5-4662-835f-af65acadd9b6 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization In International Conference on Machine Learning (ICML)
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa67e5e7-8c7a-4cbc-8cd2-f5f616578b5c · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Post-training 4-bit quantization of convolution networks for rapid-deployment
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9daf933b-ddf0-44b3-9907-a617827aab08 · outbound
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9037c0f-f0b6-4073-b2a2-f85c2c71e183 · inbound
Two-dimensional early exit optimisation of LLM inference QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.