Pith. sign in

Paper Citation Record · LEDGER

TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2404.11912.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.11912 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:28:03.971156Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae37c6f8-7817-4566-a999-0629ccfaa6a5 · inbound

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference cites this paper.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:16:32.005322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:676bbdce36d270062660f8f7ef84ede3ab517ee6a328cae2be10d7d37bc28cc4

Observation 55dec2e9-54fc-4f51-bf95-c3ff9eecc11d · inbound

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective cites this paper.

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T00:45:56.928374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:45:56.928374Z digest=sha256:60306ffa6bbce5980300c47b044096ebefb7394b5e5f6e6d8bac8c98745bb368

Observation 9ad286d0-5105-4022-8ddd-c8bd2ec549bd · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.971156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.971156Z digest=sha256:d629c804fde7c1e9e96d468f3275e0e5916dc3009f785a2c0e1881237f7c0fdc

Observation ec1d226a-ee4e-4caa-8427-5d7dc4317c79 · inbound

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding cites this paper.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.755203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.755203Z digest=sha256:7971918d176a57e3bc39c4ba61c6b1dbd5e5dfc3affcf94d22b4b089a6aaa08c

Observation b6f382d9-400e-460b-ab97-1811eeb6cb34 · inbound

CLaSp: In-Context Layer Skip for Self-Speculative Decoding cites this paper.

CLaSp: In-Context Layer Skip for Self-Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:12.564760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:12.564760Z digest=sha256:95cb81ceae8e06980ace9624ad41a21db735e45d50bcb83e18d59efce4d7b3c1

Observation 2f83e9bf-2884-4795-8cbb-e5712a61ad6e · inbound

POSS: Position Specialist Generates Better Draft for Speculative Decoding cites this paper.

POSS: Position Specialist Generates Better Draft for Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:15.669373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:15.669373Z digest=sha256:d6c53137cede2da0fcd88deb895eb9519b3d98a91857e311c0c54aab64c522b4

Observation ee44efbc-ac4b-45df-b458-0d636cf0ee21 · inbound

Rectified Sparse Attention cites this paper.

Rectified Sparse Attention TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.969030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.969030Z digest=sha256:c075cd5551cc584d427b75bc52b213adc9a91a9fbf7d92463defa32bc388b9e0

Observation db541a76-d744-4267-bcb4-3b42e9e53540 · inbound

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention cites this paper.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.855974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:53.855974Z digest=sha256:73e773f35d420acd0387b4debeafa937453bb4264afc5269554dcec77525d6b2

Observation 36684d3a-22d2-44c8-b31c-8270bb5e6e23 · inbound

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding cites this paper.

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.439354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T06:58:07.996335Z digest=sha256:f5029ca1f9a954e01d3090e28d5808590133657049b61c6387d6deff9e140969

Observation d0651bf1-832a-4558-a756-c7719002659d · inbound

DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution cites this paper.

DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:33:24.538998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:27:23.866704Z digest=sha256:75198cae892449b631bd7461020ee020903b94b1062803bd1d0464669e0c6284

Observation d390899a-2759-4015-99cf-d2c4b607c606 · inbound

BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding cites this paper.

BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:22:46.262797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:19:13.872892Z digest=sha256:56569f6245ee13598b8b2aaaf2de885bd203c5f62e6b2672f95ca5ea21f0f2ce

Observation da00e215-5224-4883-889e-0477e24aa329 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.628221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:09dcc97590ebb634c1ec79f964a3762991bb55ddbcfe7c0e57cab4523369b6c2

Observation 2937639b-229c-432b-ba4a-e9b7132bee46 · inbound

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding cites this paper.

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.999703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:11:58.636939Z digest=sha256:7910ebd55f4f9000c73b5cb332ee9db2a62f0e6633578cfa754feb1f1c8aa4bf

Observation 5558557e-37b4-413d-8ca7-cdd4310a0237 · inbound

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM cites this paper.

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:21.309531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T07:19:48.530272Z digest=sha256:519d5d489c90e9f42195928a78f36b8f518cc0180e868d8ae3b2057fd0a37ef2

Observation a50d6cf9-b693-4e7b-91d3-d5727e8a9b0a · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:42.358155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:42.358155Z digest=sha256:7513bdb7a56387a17d5f87b971b9f135fcd66298f03f7905334bbd0c3a95094d

Observation 4bcd002c-e92f-4361-9f13-78306a896d43 · inbound

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes cites this paper.

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T12:06:59.121173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:06:59.121173Z digest=sha256:af24bc1acbe1c38b4fd54b3e290413e8b7e748486f728260c9962d682c87408b

Observation 3c1aa857-587a-4a6c-975c-7ccedd3094b7 · inbound

A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding cites this paper.

A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:28:04.834122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:28:04.834122Z digest=sha256:efef69685962ebe140bbfa5516798767b8213defa75ed8412d19b033dd1b11fe