Pith. sign in

Paper Citation Record · LEDGER

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator

As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.00797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00797 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:15:08.413962Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82e96aaf-b95a-4a4a-92a5-11c2a30b477d · outbound

This paper cites GPT-4 Technical Report.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.849464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.849464Z digest=sha256:b2150a48e6da4ff7f0b9e33eaf9ccad70f201cb498cb51272d940b63201d8cc6

Observation 72b6c78e-1280-462d-a626-1b509ed01e59 · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.326845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:07.859756Z digest=sha256:d6aad580bf2b18c9829ac8e4b1a35eb40258135796b755089e87c830218429b7

Observation f0a9f903-772b-4d81-ba86-1f5b0b72b145 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.260429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:07.881035Z digest=sha256:79b5c90a170e1c645b9ee43dbcbd4dc709df55859299c5f70259712ec815edc8

Observation 1f6b3c51-5858-4aa2-aa87-8a6e3bdecb84 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.905855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.905855Z digest=sha256:727a42af8bb5e6b9ea127a0b80ab0c5737af6461270da7502144783f68aca733

Observation af994ca5-914f-436a-ba39-997971429553 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.925788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.925788Z digest=sha256:7da5f4671063c0525a52848c757289348e29b699921ff547bdeeeee332ffa810

Observation a266e578-3faf-4682-8594-a64aaca11cd0 · outbound

This paper cites Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:07.936932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:07.936932Z digest=sha256:c34c92d3867adb15deef621ff36065fd233392bfdea24ef0f8e20fbc5d2d9d07

Observation bce1e27a-5595-499d-b96d-16313904e1ae · outbound

This paper cites Ramulator: A fast and extensible dram simulator,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Ramulator: A fast and extensible dram simulator,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.122654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:07.946905Z digest=sha256:f2cb29ad491a407eb019b35215dc1de983304f7fe31590f890cec155961d6b47

Observation c13cd7a9-0e80-45d4-a83b-1f4ead0f21e0 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:09.045540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:07.979589Z digest=sha256:83ba0f249116dd0fc5d0c0ca1fb458202537835c5be8f1bc0d7e3f6f45d424b9

Observation 8c532b4d-d698-4c75-91b7-32fba6c84da2 · outbound

This paper cites Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.952344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:07.998054Z digest=sha256:57fcc3bf37e1864becba4b9693ce0bad519673800c6e95ba66fbfa22bd7cc7b9

Observation 17c0dbb8-0ca1-4db4-ac07-30d9a18fb3bc · outbound

This paper cites Online normalizer calculation for softmax.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Online normalizer calculation for softmax

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.020611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.020611Z digest=sha256:699023592cbb4eb5104d6ce8a927cbeb58baa5b06d2e8bc0210ec1a25549db04

Observation 2727282a-607f-4d4a-ad12-ec698b5a3a26 · outbound

This paper cites Cacti 6.0: A tool to model large caches,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cacti 6.0: A tool to model large caches,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.845873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:08.049142Z digest=sha256:c4b4ecd76d0c58b689c2a4b72972f3dd9f6d17ad00045db262324379fedfc85a

Observation cfb3fef2-0844-4581-9a1c-100168a196e5 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Compressive Transformers for Long-Range Sequence Modelling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.085353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.085353Z digest=sha256:ab32a7edb441deddec07ed2c694be051a6ee8daad1163cf6d8edf0d1acdb0276

Observation 4613ce40-fd13-4dfe-b3a6-1785a1e4d68f · outbound

This paper cites Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.788210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:08.129401Z digest=sha256:a7ac12a703489af4cd98fcc2635d7421a3ac3af6d97884a192ed0806632664fc

Observation 695b8885-89e2-48f5-ae99-a1b45604d5f1 · outbound

This paper cites Softermax: Hardware/software co-design of an efficient softmax for transformers,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Softermax: Hardware/software co-design of an efficient softmax for transformers,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.166263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.166263Z digest=sha256:2b8aa7ffe8e668ea74fc87ed2afab5dcba6a61bc59fe059bf1123e9af8fdc49d

Observation 39010cac-1c04-4604-bd31-4cd13cf2be2a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.202870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.202870Z digest=sha256:e3b76e766193d5437566099ecfba83bcc72757b3f23428ab214232ee368fd7d2

Observation 06dcc06d-a9e5-492e-b322-38fc81ecfe8a · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.237183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.237183Z digest=sha256:6023860de31e40cb1ddf3413ec44e629abfc4c43c540b415ecd135cfd4f1dab1

Observation 77954335-5988-4f1b-bf7e-71adab422002 · outbound

This paper cites Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.695573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:08.274907Z digest=sha256:f20e1ae7ebaa6995d110e57fc7c8d96ec1c8a1042ddcfe14c7b68e20b4597d69

Observation 6f659055-ae9e-4c17-9b50-cecc7f8b9193 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Efficient Streaming Language Models with Attention Sinks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.309519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.309519Z digest=sha256:226423068fbeb3e087f96a191e816fd9f18f05d4fe1a24837306df42aceffe12

Observation 7a12d836-3440-47af-b3b7-859070fee557 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Orca: A distributed serving system for {Transformer-Based} generative models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:15:08.638398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:15:08.345273Z digest=sha256:a74c5dd0570a09d87af2128bf261dd2e5115e689e6e0e139fe61070e2d02769d

Observation 4e5f3e1d-e7ca-41e7-b216-ccac6a4f76db · outbound

This paper cites NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.374201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.374201Z digest=sha256:0b626b6a6df83de4e30ba2467c5a8afb802c84fb0089a84ab6b659a599dbce6d

Observation 9d67cf50-a008-43ef-8f0b-22996fae256b · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:15:08.413962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:15:08.413962Z digest=sha256:683b584c965d0f8aa4225da729f4eb33616defd032728860466fd60fc2230269

Pith citing papers

No inbound Pith citation observations are available.