Pith. sign in

Paper Citation Record · LEDGER

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization

As of 11 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2506.22396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22396 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:23.901164Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-14T22:59:07.065193Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T22:59:34.279645Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbdea40f-a78e-4bf2-a233-45def0e3706a · outbound

This paper cites This ": processed all layers.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This ": processed all layers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:29.345819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:21.644110Z digest=sha256:b1061711c0d85088e0263a455b0aa9caa04b9f9e1da954ed30159dc7a2d075ee

Observation 55d60984-3f8a-4332-ad5b-83b8aa8f4132 · outbound

This paper cites this ": KV diff 1.00 < 0.30 -> Write.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization this ": KV diff 1.00 < 0.30 -> Write

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:29.210914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:21.711859Z digest=sha256:8b000780fd88f1c0fd7607f9d678b20373f44610fd70bfbd2747db2fe10670fb

Observation ceffd707-cb7d-43bb-a4ab-7af4522662a4 · outbound

This paper cites This " +.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This " +

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:29.042472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:21.813508Z digest=sha256:041a89c77d2cb2cc983e57b604f6bb8a6764052af91ac5c33ce61f81a2f39ae6

Observation f923ebb3-9e92-4cd0-9245-b8626840ec47 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.844437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.844437Z digest=sha256:c6322e1da2812daa28c082628ea98249d6d9651bb33cd644d0c1d25ce04ba615

Observation 7fc147b7-69e6-4816-baf2-79182af41f32 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.927295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.927295Z digest=sha256:94947f21543075f591242b8baf5b16908a147587ae288da9d834009bfbec36cc

Observation d563ab34-0ddb-4def-8e66-9e57cdad4304 · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Adaptive Computation Time for Recurrent Neural Networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.989489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.989489Z digest=sha256:372bbfd4719aab8825ee46936cd3e3ff85d00b0ce389dfb517d3cc03cc80ab15

Observation 10675910-233b-4717-b9d8-d1a0a413bb93 · outbound

This paper cites Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.184745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.184745Z digest=sha256:48c5e85a81b973c8c9a68b7ba664faaeba5f5e559da42aaf6f2804719a9e2488

Observation 6e1305a7-c7ab-4ee0-80cc-4a5aeb0c3cb4 · outbound

This paper cites OpenAI Blog.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization OpenAI Blog

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:29.706660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:21.302545Z digest=sha256:c5793f50dd6075a2414485037476c14924c699e8db877b87b40063349c89f5ac

Observation 9830d11b-602f-42ed-966e-906875262a7b · outbound

This paper cites Transactions of the Association for Computational Linguistics, 8:842–866.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Transactions of the Association for Computational Linguistics, 8:842–866

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:29.525120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:21.421070Z digest=sha256:1e1cec2df8e01803338d002817b05400c7be7ba69b4b67295f69d2b18d13a29c

Observation 3005d7e8-64b6-4768-9db1-f58c93b9e6b5 · outbound

This paper cites and ": entropy 0.23 -> 2 - bit quant Token.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization and ": entropy 0.23 -> 2 - bit quant Token

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:28.888395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:21.931048Z digest=sha256:3019905ce88292c566b2bc264879a9569eb53c0915f1924527ca5c0b4b07dd0e

Observation 5de69b14-c194-45c1-8a33-36faf45a01f1 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:28.721423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.027589Z digest=sha256:6bfbaba5842d827b7ca919bddbc103c7dff3f5aa27060703e86a450754d2537b

Observation 70b860aa-9582-4900-a298-c52495bfbb84 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:28.464567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.088279Z digest=sha256:d7b163ae04522211281d7e54741509f16c95b9b608d29aecdf8dace360b05c67

Observation be1bc39a-c227-454f-9c9b-023fc1d7c74b · outbound

This paper cites the, ” “of,.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization the, ” “of,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:28.130054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.152171Z digest=sha256:36e38ec31af97e184666d43654200c5bab87fbe67dc983907e6fe5f84e2e0dde

Observation 340a6210-5183-4572-bc9e-972aa58128c8 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:27.797728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.225643Z digest=sha256:943f9f322675b4669c1e7f57d6d5757248891d38f3cb4e0a0654afbd2abe47eb

Observation 8e967772-e3fd-432c-971b-ea39e29bc34a · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:27.481988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.308022Z digest=sha256:8169de3b0c499f0caf2219ede3628afbf5c4fd03991d2e656a972e44b0d21c1a

Observation b44abbd4-911d-4b2d-8854-64d7dd03a1fb · outbound

This paper cites skipped after Layer 2.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization skipped after Layer 2

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:27.259969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.378521Z digest=sha256:25f780ebc951155f6850f42889748769da034e0767acc0980087f35f61a70832

Observation ac9c5c76-258e-483a-9422-72051c13e320 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:26.990889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.475657Z digest=sha256:4bdb90963d20ebca323757a77aab9dc4f8a5fb05c69e160fdb8bcbedb861c83d

Observation 50b1b291-d620-430c-afe5-a794443dc6fb · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:26.661016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.548011Z digest=sha256:8b0acf16714693b5a48aff5cf3c81440484fca8cb243d75e5b6f1fb36d954832

Observation ff59b215-10fa-4728-a5df-f16c00033a54 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:26.307801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.634481Z digest=sha256:ab18634e3f1a2b1c4ece82bb5cfb1b09d16e94d9e74370578e91e1768a439209

Observation 320abb27-1600-4517-9b8f-8980131a5ebe · outbound

This paper cites This conservative threshold ensures fusion only when representational collapse is semantically safe.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This conservative threshold ensures fusion only when representational collapse is semantically safe

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:26.002069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.724335Z digest=sha256:36a1d1e49b151898e1c465b3b3ce56802454ac891cb1afc615fc2eca59b181d8

Observation 13489094-6c33-4013-9521-78b40884a72b · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:25.699415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.804178Z digest=sha256:e3b984218ec516cdbfb2c0ddde0e04f7d269a9c6b372a9f7c31323e78e201a57

Observation a3a64abf-0ac8-46a2-bd14-9a857917d78e · outbound

This paper cites We choose the pair (τlow, τhigh) = (0.3, 0.6) that achieves a strong Pareto frontier: ∼ 8.6% additional FLOP reduction with less than 0.1 perplexity change.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization We choose the pair (τlow, τhigh) = (0.3, 0.6) that achieves a strong Pareto frontier: ∼ 8.6% additional FLOP reduction with less than 0.1 perplexity change

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.606207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.902853Z digest=sha256:8590bc81d5c1796b469a910498ac79c62f29f421c0520b181fda4721f4c3ccae

Observation dff6614e-5875-45ae-b55d-43a4879dbd19 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:25.456484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:22.970994Z digest=sha256:a96e9cde7908b53810dcb435334190eadb8224e23145c244a96d4d201e6ffc57

Observation 9b1963e1-8a94-4e7c-b5d8-77fe94faa165 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:25.334389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.067941Z digest=sha256:94a138e4175a200a38cb7dd1703d096cba08515f84fb31a6d261f83596112207

Observation 5780c87f-a0be-4ef7-aac8-637ceecbbc95 · outbound

This paper cites Pareto Surface Analysis.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Pareto Surface Analysis

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:25.204098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.157049Z digest=sha256:a167bf409ba72e7fefd2c4d6117677701261c334795f6a783d326f82cede9fa3

Observation 2632d43f-6fdb-40dc-a5a3-e86ccce7e131 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:25.073338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.292848Z digest=sha256:a87e4fad596620a31fb15bfcd1700d1d03375b3f12d829922eb715632c0228c3

Observation 5a48624b-df50-437f-a210-2fa41d532593 · outbound

This paper cites an unresolved cited work.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:10:24.929281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.410599Z digest=sha256:598e3fa57be132abd9a97be4706e710e678cd990128783d22f4d36ccc7ade596

Observation 109bdff8-33a9-4348-a66a-093ce5cc00fb · outbound

This paper cites We hope this work encourages the community to adopt tools like CodeCarbon not as post-hoc profil- ers, but as first-class citizens in the deployment pipeline.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization We hope this work encourages the community to adopt tools like CodeCarbon not as post-hoc profil- ers, but as first-class citizens in the deployment pipeline

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.813090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.502380Z digest=sha256:91bcef0ea95b1a5d18bb5738b7a5231608f3099b96ef155354bfe4bf2180bc38

Observation 21bf7a74-74a1-40a4-88b5-0dc614ac3f85 · outbound

This paper cites I.2 holds, halt the token unconditionally.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization I.2 holds, halt the token unconditionally

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.715856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.588970Z digest=sha256:fd502471cba3ec382bb9073c7a7a41fb6a65212f3b410ca204f0fc03e6859b8b

Observation 7f17e496-1fa0-4250-a091-099349017816 · outbound

This paper cites I.3 holds for any u, merge (t, u).

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization I.3 holds for any u, merge (t, u)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.575835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.675410Z digest=sha256:b6e54a4efc0aece082d53fe911e00f8f5fab87ebf4231b3d8b150f8ba90834d0

Observation 6a014c21-1968-4921-a1bf-73f7835c4458 · outbound

This paper cites This priority is grounded in halting the provision of computational savings without representational loss, while fusion entails approximation.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization This priority is grounded in halting the provision of computational savings without representational loss, while fusion entails approximation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.483090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.769522Z digest=sha256:99e5f9023cad724993c4661dceb34068e06d6ec1379e126c59b85aa390a616aa

Observation a247a1bb-80cf-4184-a20b-6a4175ef868f · outbound

This paper cites the”, “of.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization the”, “of

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.333934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.831764Z digest=sha256:aa6b93de2cbf804cec472e7dc2732c79f517e8cacfe59861fef6672172b7f62b

Observation 53f694ef-434c-44fb-8cba-210f44d7265b · outbound

This paper cites tokens.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization tokens

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:24.153732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:23.901164Z digest=sha256:f9eb5efe9d094d238d4ae5a01b5767ed2ed8fbda177bff326f09599b8c909633

Observation 049f2228-6a3f-4ccf-86d7-8905790b5e24 · outbound

This paper cites Shaojie Bai, J Zico Kolter, and Vladlen Koltun.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Shaojie Bai, J Zico Kolter, and Vladlen Koltun

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:29.876336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:20.558635Z digest=sha256:1f73b0c9611f97267f11f8f1e1d52c1052d7ae12787b3a3e4f139aa4e9710aef

Observation 76f29783-9953-4308-a35a-b0f1be593238 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Fast Inference from Transformers via Speculative Decoding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.087718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.087718Z digest=sha256:6e9802ab91ebb591ddc88356e8ed583c21f742443045d7e5ae240ad4e56b999d

Observation f82ae740-22a5-4662-835f-af65acadd9b6 · outbound

This paper cites In International Conference on Machine Learning (ICML).

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization In International Conference on Machine Learning (ICML)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:30.034930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T22:10:20.466955Z digest=sha256:00e9117b6f44a0458d94f149dd04d23f0b5e549ee49a590dccc666d741111441

Observation aa67e5e7-8c7a-4cbc-8cd2-f5f616578b5c · outbound

This paper cites Post-training 4-bit quantization of convolution networks for rapid-deployment.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Post-training 4-bit quantization of convolution networks for rapid-deployment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:20.694641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:20.694641Z digest=sha256:ed05ab46f7a2154e77cfac0bba434f7340efba534727cac0c89e8897214e5713

Observation 9daf933b-ddf0-44b3-9907-a617827aab08 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:21.546065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:21.546065Z digest=sha256:bc7f8e1188b031f5c79637e3af95a59ef62ee3cd30495fbba0c9686ef78d5d5e

Pith citing papers

Observation c9037c0f-f0b6-4073-b2a2-f85c2c71e183 · inbound

Two-dimensional early exit optimisation of LLM inference cites this paper.

Two-dimensional early exit optimisation of LLM inference QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:59:34.284901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T22:59:07.065193Z digest=sha256:5c82a35d959d8ca8aa92d50f2fb89ff4b1689e966c0c1fcb44623c7e6110c76f