Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:15:08.413962Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.00797.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:15:08.413962Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 82e96aaf-b95a-4a4a-92a5-11c2a30b477d · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b6c78e-1280-462d-a626-1b509ed01e59 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0a9f903-772b-4d81-ba86-1f5b0b72b145 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f6b3c51-5858-4aa2-aa87-8a6e3bdecb84 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af994ca5-914f-436a-ba39-997971429553 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a266e578-3faf-4682-8594-a64aaca11cd0 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Aˆ 3: Accelerating attention mechanisms in neural networks with approximation,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce1e27a-5595-499d-b96d-16313904e1ae · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Ramulator: A fast and extensible dram simulator,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c13cd7a9-0e80-45d4-a83b-1f4ead0f21e0 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c532b4d-d698-4c75-91b7-32fba6c84da2 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 17c0dbb8-0ca1-4db4-ac07-30d9a18fb3bc · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Online normalizer calculation for softmax
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2727282a-607f-4d4a-ad12-ec698b5a3a26 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cacti 6.0: A tool to model large caches,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfb3fef2-0844-4581-9a1c-100168a196e5 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Compressive Transformers for Long-Range Sequence Modelling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4613ce40-fd13-4dfe-b3a6-1785a1e4d68f · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Deepscaletool: A tool for the accurate esti- mation of technology scaling in the deep-submicron era,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 695b8885-89e2-48f5-ae99-a1b45604d5f1 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Softermax: Hardware/software co-design of an efficient softmax for transformers,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39010cac-1c04-4604-bd31-4cd13cf2be2a · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06dcc06d-a9e5-492e-b322-38fc81ecfe8a · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Spatten: Efficient sparse attention architecture with cascade token and head pruning,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77954335-5988-4f1b-bf7e-71adab422002 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Cosa: Co-operative systolic arrays for multi-head attention mechanism in neural network using hybrid data reuse and fusion methodologies,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f659055-ae9e-4c17-9b50-cecc7f8b9193 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Efficient Streaming Language Models with Attention Sinks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a12d836-3440-47af-b3b7-859070fee557 · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator Orca: A distributed serving system for {Transformer-Based} generative models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e5f3e1d-e7ca-41e7-b216-ccac6a4f76db · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d67cf50-a008-43ef-8f0b-22996fae256b · outbound
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.