Pith. sign in

Paper Citation Record · LEDGER

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2606.07713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07713 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:32:56.444150Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:08:09.010538Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cce840d-df48-4bb0-a082-325f5c2d61b8 · outbound

This paper cites FATHOM: Fast attention through optimizing memory.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FATHOM: Fast attention through optimizing memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:856ebb6fe15a6aa3dcab9c29e5ed66e42c38379131e1ef27218499596a7d4bb2

Observation a35b2311-1aa3-49ac-a113-eb678af4a77a · outbound

This paper cites MOSA: Matrix optimized self-attention hardware accelerator for mobile device.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels MOSA: Matrix optimized self-attention hardware accelerator for mobile device

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:b98adb880d6c891f62b059f6d447cad451f17256b67f3e69bd7fb2a4695147e8

Observation 2ff700ae-f76d-4b9c-b841-57e88d6cbaff · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:37:09.301147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:069d79921da9420ca1fc135dafe2c8375391f0dce8cc9c0902f52ffe11ffe8bd

Observation 9e6bbb37-4863-4ed9-a3e4-4f72e56f162b · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:37:09.295708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:19e99a2278a506e70f406ee65a5968bcadff935fb5c8e03d33e8ec230f63ce0e

Observation 6605622a-6c48-4fd0-b93a-e5e1ab789318 · outbound

This paper cites Hardware considerations for tensor implementation and analysis using the field programmable gate array.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Hardware considerations for tensor implementation and analysis using the field programmable gate array

Reference 5

Resolution
verified exact
doi, observed 2026-06-27T22:41:23.898840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:6b3daf626c4c755e31d490791133975c73a2363cdbc3c36214fe5121a11baa07

Observation bc413ab4-7cbb-4e9f-86cf-7049052967c2 · outbound

This paper cites Realizing mathematics of arrays operations as custom architecture hardware-software co-design solutions.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Realizing mathematics of arrays operations as custom architecture hardware-software co-design solutions

Reference 6

Resolution
verified exact
doi, observed 2026-06-27T22:41:23.900855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:1b4e1317707f51ddc26605fb048f3a4a05401b357a64556e2389160b818be3a9

Observation 16216584-1b2d-4b85-8a07-c0fa99c83335 · outbound

This paper cites Processing in memory for mathematics of arrays operations.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Processing in memory for mathematics of arrays operations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:b4d5a70d48e41ba04c6162d5b86def55bb7906fe2e02543334bdc28b3a900e52

Observation 892184b8-f20f-4656-958a-3a9e40760422 · outbound

This paper cites A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels A Fast Optimization View: Reformulating Single Layer Attention in LLM Based on Tensor and SVM Trick, and Solving It in Matrix Multiplication Time

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.298326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:b267ed019c770239cdc9a11865e8a0e2f40805a93e9d098e48611e411626403e

Observation 1a516b0c-d999-4a5d-a791-1a5bb6b91bc9 · outbound

This paper cites an unresolved cited work.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:756971bce257d0e49223f2a9c8a198fd6d890dd5aad1a7b5c7316dd5378017fc

Observation d1055bb5-dbff-4801-914b-ad13cf3a9eb9 · outbound

This paper cites A ^3 : Accelerating attention mechanisms in neural networks with approximation.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels A ^3 : Accelerating attention mechanisms in neural networks with approximation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:ffaec55468113ec80c6145a204ec8291f64a25360d3ca36b990841845f33c852

Observation 318cae26-df65-49f6-b0e8-cce56e9e9e3a · outbound

This paper cites Acceleration of fully connected layers on FPGA using the Strassen matrix multiplication.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Acceleration of fully connected layers on FPGA using the Strassen matrix multiplication

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:895f25abccc6de32965b9fcd8d7a66601708e5d2ccf3c5a78556d56c38fccaf5

Observation 1b2f47cc-08a3-4c42-acf9-80b3ac870984 · outbound

This paper cites Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.293187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:c79f78a7435c3d825c77846feba86c79a6c118cfc18e37f43d83da6b6ff5fc62

Observation f7372ef7-b175-4c53-93e5-2bf4e08abd07 · outbound

This paper cites Research on matrix multiplication optimization and deployment method for heterogeneous platforms.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Research on matrix multiplication optimization and deployment method for heterogeneous platforms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:28a726e08c012fee88c65c08ab251833b3741ede38dd31c7b872278d60a82be2

Observation 2a5736dd-15d8-49d6-86c5-9d5b4fb35d16 · outbound

This paper cites Array access and performance regarding numerical algorithms.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Array access and performance regarding numerical algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:0cd91aa72aeb88bf4b2ff467971f40b4b8ad4daafdae9f8019f8967a347416c8

Observation e41d9201-4449-4f98-9682-a16cd1455ac3 · outbound

This paper cites A Mathematics of Arrays.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels A Mathematics of Arrays

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:d254b8cd6fbafde33b0313ad8d0a45ff902f36f61a318be808b9198fb342b470

Observation de69877b-f1c1-409c-9965-6f0457587a40 · outbound

This paper cites From array algebra to energy efficiency on GPUs: Data and hardware shapes with dimension-lifting to optimize memory-processor layouts.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels From array algebra to energy efficiency on GPUs: Data and hardware shapes with dimension-lifting to optimize memory-processor layouts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:37:09.280982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:bfa5f0b79c20948754e51d07e2023aacea6ee6826235c76c85291348ea2e3ae3

Observation dc789cdc-e395-4774-b16e-ecbd46091da6 · outbound

This paper cites Towards automatic, predictable and high-performance parallel code generation.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Towards automatic, predictable and high-performance parallel code generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:23d9bf80a271ded2ceb31382675833196b24f8e5d0591c8748e6fd71b721ee1c

Observation 82f8b2be-9d72-4e49-927f-43f4e20e1efe · outbound

This paper cites New mathematics for computer performance: array algebra and cost functions.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels New mathematics for computer performance: array algebra and cost functions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:c6a2cfbbaf962dd031f43977de9572c0a1b4dab76a4e4c2422474f7bdf449dd0

Observation 288b6404-5124-4727-8712-cc421e9ab0a0 · outbound

This paper cites OPTIMUS: Optimized matrix multiplication structure for transformer neural network accelerator.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels OPTIMUS: Optimized matrix multiplication structure for transformer neural network accelerator

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:5e0b0c535fa37de34d0725740349f8ab35c9dd8c4a8b26ac24c17fb0ee2bcabe

Observation 411e3766-809e-4e95-b3ab-8e7a5647c639 · outbound

This paper cites FACT: FFN-attention co-optimized transformer architecture with eager correlation prediction.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels FACT: FFN-attention co-optimized transformer architecture with eager correlation prediction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:8b231ddbc82ef316c10a60910361ca0c9d93a55f6cfb4823affbbf298bb16b07

Observation ebbeb976-a8bf-4c79-a53a-404cc2581dbf · outbound

This paper cites High-performance Gemmini-based matrix multiplication accelerator for deep learning workloads.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels High-performance Gemmini-based matrix multiplication accelerator for deep learning workloads

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:59a2a4ff30a635c2f08812798d20ec4f533af46b1dc60a805d1f074b813bbc0c

Observation b09d5781-1cb5-4a9b-9b76-5dc65048c819 · outbound

This paper cites Design and implementation of a BRAM-banked double-buffered matrix multiplication accelerator for transformer models on edge FPGAs.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Design and implementation of a BRAM-banked double-buffered matrix multiplication accelerator for transformer models on edge FPGAs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:ade5f1467322e1e7a4cd13de8e534e823e017ac17163c05007b97ceab1b173be

Observation b706a879-9be6-464d-a411-9e548f1f4de6 · outbound

This paper cites Improving the performance of DGEMM with MoA and cache-blocking.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Improving the performance of DGEMM with MoA and cache-blocking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:17f5e8c7a2d37d5b69479ecd3541d74fb311b5b117ee9101263ef4c05991c8d3

Observation 2cedba33-dad1-4d9a-9f92-2e71ebd5d115 · outbound

This paper cites Threaded multicore GEMM with MoA and cache-blocking.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Threaded multicore GEMM with MoA and cache-blocking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:2e2596805897a1f023de95107e03caac7740841fb89238d687f2d6ab24060613

Observation 01cf1fa2-2842-4610-afb9-c30081b9588a · outbound

This paper cites Attention is all you need.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Attention is all you need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:a1d9a1287d951ba267e99a1dbb8b4bb736fd556f498123bc167f5666f03e2196

Observation 357f9832-6c53-4e8a-ad71-fc104d91e53f · outbound

This paper cites Hardware friendly transformer optimization with dynamic attention matrix fusion.

Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels Hardware friendly transformer optimization with dynamic attention matrix fusion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T22:32:56.444150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T22:32:56.444150Z digest=sha256:9728d3a195ae91248e108a00e0f95b95b8dee4dfc35f2abc16003688dd70e08e

Pith citing papers

Observation 487c3671-2490-403b-b5b6-ecbcad1095be · inbound

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel cites this paper.

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T13:08:09.010538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:08:09.010538Z digest=sha256:2cd62295a4b47c10e84acfcb6cff1622c2ddbb5edb605a1ab757db25f4f2863e