Pith. sign in

Paper Citation Record · LEDGER

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.00797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00797 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:15:08.413962Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82e96aaf-b95a-4a4a-92a5-11c2a30b477d · outbound

This paper cites GPT-4 Technical Report.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.849464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.849464Z digest=sha256:b44638501df35d3d6d3946551a0d362ba3d56d9ce017e0d548d7d58600a979af

Observation 72b6c78e-1280-462d-a626-1b509ed01e59 · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.326845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:07.859756Z digest=sha256:c6e12deec970e1684da19cf830ddfdf73e5634a60f4fbb3b25c940161f7803a4

Observation f0a9f903-772b-4d81-ba86-1f5b0b72b145 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.260429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:07.881035Z digest=sha256:78806b837b797f953828582e2aeb3a1a31dd288d86eb934d596602f7706c0ad5

Observation 1f6b3c51-5858-4aa2-aa87-8a6e3bdecb84 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.905855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.905855Z digest=sha256:7a260cdba78d25876af4cee68cc7176f8546fc529e8bf42de5d402f25ca07a5e

Observation af994ca5-914f-436a-ba39-997971429553 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.925788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.925788Z digest=sha256:69e42e9f7266940215b1d74993708111836cd365e9e602ebfc38ab5e56215dc5

Observation a266e578-3faf-4682-8594-a64aaca11cd0 · outbound

This paper cites Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.936932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.936932Z digest=sha256:58eeaae49bc6b98048cf4f3341a1ce1f672a171ceb8bc758592cf806230bb841

Observation bce1e27a-5595-499d-b96d-16313904e1ae · outbound

This paper cites Ramulator: A fast and extensible dram simulator,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Ramulator: A fast and extensible dram simulator,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.122654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:07.946905Z digest=sha256:4f1bcb52b959d1b539ed29375360ea5702078265e9b4f8468202a7c76d391eb6

Observation c13cd7a9-0e80-45d4-a83b-1f4ead0f21e0 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.045540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:07.979589Z digest=sha256:aef032b466cf0fc95f00ea195d2da442885e471dae59711257c0da3e35cd90d8

Observation 8c532b4d-d698-4c75-91b7-32fba6c84da2 · outbound

This paper cites Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.952344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:07.998054Z digest=sha256:50d306d689197fa06bf05309a66c30ac4b1385c1a348ed03d90c2045526e14ce

Observation 17c0dbb8-0ca1-4db4-ac07-30d9a18fb3bc · outbound

This paper cites Online normalizer calculation for softmax.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Online normalizer calculation for softmax

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.020611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.020611Z digest=sha256:80082722fd9b34bce3aec0125214017378abc71fe5d8aec47cd4d19b7d533122

Observation 2727282a-607f-4d4a-ad12-ec698b5a3a26 · outbound

This paper cites Cacti 6.0: A tool to model large caches,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cacti 6.0: A tool to model large caches,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.845873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:08.049142Z digest=sha256:e3e999a268ed19f1e256fb1efe07ef241346704cdfd5434a6ea0c51031f1b9e2

Observation cfb3fef2-0844-4581-9a1c-100168a196e5 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Compressive Transformers for Long-Range Sequence Modelling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.085353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.085353Z digest=sha256:2a35a21516741b4bcda09a62d85e011ebc0e1f7ef3b08bc4ce180927fec599a5

Observation 4613ce40-fd13-4dfe-b3a6-1785a1e4d68f · outbound

This paper cites Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.788210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:08.129401Z digest=sha256:178b2516cd32ea8105bf5f297524ea52d5bb97af9d91128d5c9fab323e8dc531

Observation 695b8885-89e2-48f5-ae99-a1b45604d5f1 · outbound

This paper cites Softermax: Hardware/software co-design of an efficient softmax for transformers,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Softermax: Hardware/software co-design of an efficient softmax for transformers,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.166263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.166263Z digest=sha256:367b31c6881204b14dfcbbedec1e18576b849270cc459453285b7908f8a49987

Observation 39010cac-1c04-4604-bd31-4cd13cf2be2a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.202870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.202870Z digest=sha256:498d143cf991e3cbd9acae5f57df29bf874d57abe02e34d77894fdfa5f3d76de

Observation 06dcc06d-a9e5-492e-b322-38fc81ecfe8a · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.237183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.237183Z digest=sha256:1586492ea3d092c0d4333eed5748f379f1593efc562dd08d5605d13c6cae7367

Observation 77954335-5988-4f1b-bf7e-71adab422002 · outbound

This paper cites Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.695573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:08.274907Z digest=sha256:a094fd7594486ff60228fa158b99b5fa300e05c63a71e13b1c31bcbd504129cb

Observation 6f659055-ae9e-4c17-9b50-cecc7f8b9193 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Efficient Streaming Language Models with Attention Sinks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.309519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.309519Z digest=sha256:842e0c41fd8f90b8761e95d5ade381c89d3c108eeaf9b656ff8c9da82ae1db17

Observation 7a12d836-3440-47af-b3b7-859070fee557 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Orca: A distributed serving system for {Transformer-Based} generative models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.638398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:15:08.345273Z digest=sha256:72bb5709d1c3c3631ad6e5fcd5e6c30e417f940fd1800c4646412af994a60f61

Observation 4e5f3e1d-e7ca-41e7-b216-ccac6a4f76db · outbound

This paper cites NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.374201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.374201Z digest=sha256:af8ed398a7fff5725a35293f1e0e9cb0b498c24366738c7a31537e8358576ab7

Observation 9d67cf50-a008-43ef-8f0b-22996fae256b · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.413962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.413962Z digest=sha256:8fdd6197d748058218b2a9dc3b4d8cbbd5fba6494b80c03e0be40bb5505714e2

Pith citing papers

No inbound Pith citation observations are available.