Pith. sign in

Paper Citation Record · LEDGER

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.18824.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18824 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:29.161352Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6baf57a8-36f6-4226-92a5-69aa9eaf98e4 · outbound

This paper cites On the computational complexity of self-attention,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators On the computational complexity of self-attention,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.233199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:26.197619Z digest=sha256:aafd666bf4068c480e568112de4a3acfbea83e965226898653b6149240344437

Observation 2da04287-b6d6-4094-9ce7-7ba8329aa926 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.256211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.256211Z digest=sha256:89dadade441c160282b60c60bad5e8392183f3576429f59c078805f696231fba

Observation b0cfb9b3-f0e9-4eff-95a5-e7f8e75b694e · outbound

This paper cites Data movement is all you need: A case study on optimizing transformers,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Data movement is all you need: A case study on optimizing transformers,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.117193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:26.366113Z digest=sha256:985065b71dca147d507f14272f244c2df5f303ac00bcfc19c49ea10ddded490b

Observation b9234b8b-47e6-4fe6-9478-77b618a41694 · outbound

This paper cites FlashAttention: Fast and memory-efficient exact attention with IO-awareness,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention: Fast and memory-efficient exact attention with IO-awareness,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.965439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:26.504090Z digest=sha256:c24c9cb9559d0255962d39216f616af6bfa4081e15e2c366c4fe1710bfd40cc7

Observation 4b1303b3-74ce-4003-a6e8-97045b74dd87 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention-2: Faster attention with better parallelism and work partitioning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.631529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.631529Z digest=sha256:413a323aa618ed948fcd198ba2c6a1a55155d4d2c87a23a865a2de1c8309a9a9

Observation ff115aa5-8c1a-4601-a0ef-4a94f2508275 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.731577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.731577Z digest=sha256:39f2ada3af1b8d5261f944a031ecc4a1da2f5146e8848b41959d31394734e02c

Observation ead9f0f6-6a3c-4d01-8439-511de6f2d9fe · outbound

This paper cites SambaNova SN40L: Scaling the AI memory wall with dataflow and composition of experts,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators SambaNova SN40L: Scaling the AI memory wall with dataflow and composition of experts,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.841153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:26.876797Z digest=sha256:0002e6e8a94e44cda61fcf9d8b94196cb04e1a81019b4e1bfceb5e913e4063bd

Observation 555d3de4-b705-4e0b-9c89-cd24aeb2e39d · outbound

This paper cites Wafer-scale AI: GPU impossible performance,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Wafer-scale AI: GPU impossible performance,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.637534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.012739Z digest=sha256:35041384cf21354333b63dfb0af6148e2ec274c3edf589b892998992544c6a61

Observation 3ea76046-4229-47bc-9ca7-5f44509bda34 · outbound

This paper cites Blackhole & TT-Metalium: The standalone AI computer and its programming model,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Blackhole & TT-Metalium: The standalone AI computer and its programming model,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.468948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.140730Z digest=sha256:e952fdde15e6e035e409f1f8fcaa20d75eea55194689713384e1ad6faa53a073

Observation f5bedae7-44b5-429c-85de-efbe08a69ee3 · outbound

This paper cites 16.2 rngd: A 5nm tensor-contraction processor for power-efficient inference on large language models,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators 16.2 rngd: A 5nm tensor-contraction processor for power-efficient inference on large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.303301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.260624Z digest=sha256:9b3f623eee7fc2f8bc5fbdd008feaa27db33c6747b7f5900556762b648b61494

Observation 47282922-b877-4fad-93a0-853e5b5f66bb · outbound

This paper cites Attention in SRAM on Tenstorrent Grayskull.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Attention in SRAM on Tenstorrent Grayskull

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:27.394330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:27.394330Z digest=sha256:0f24af49ea4ba5f03381747ef8c78b0eb221bdb3ed7fa240ef32908854df0d9e

Observation a08011de-6c7d-4a4d-834e-bc71041f9341 · outbound

This paper cites FLAT: An optimized dataflow for mitigating attention bottlenecks,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FLAT: An optimized dataflow for mitigating attention bottlenecks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.164993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.554302Z digest=sha256:32963fd261242d10576e5ed4646d885482224aa090dd09bf00ced1d11b0fa354

Observation 2bff695c-0266-4cd1-ad94-120f6114f8ee · outbound

This paper cites Fusemax: Leveraging extended einsums to optimize attention accelerator design,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Fusemax: Leveraging extended einsums to optimize attention accelerator design,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:33.034879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.667462Z digest=sha256:1c8f82b94e4ae054525530e08ef2397aac371d9f65f45abdf612b7a54e1d5aea

Observation bb749651-27e1-4ec3-ba8e-74bbc933b1ae · outbound

This paper cites Gemini: Mapping and architecture co-exploration for large- scale DNN chiplet accelerators,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Gemini: Mapping and architecture co-exploration for large- scale DNN chiplet accelerators,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.881056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.784853Z digest=sha256:eb228258775d5ce14cccf87267b1536fc0248769e6e79a4ae0986c3a44911e3e

Observation b3bf5695-480a-476c-99a9-b3a58ff851e5 · outbound

This paper cites DOJO: The microarchitecture of Tesla’s exa-scale computer,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators DOJO: The microarchitecture of Tesla’s exa-scale computer,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.650726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:27.883144Z digest=sha256:4f27fc65c40675bf9335385beeaf83108dafc9e36f848a2a9a0ba341a0e67f54

Observation 05faa99a-a993-4ae1-af69-9b1534aa2b64 · outbound

This paper cites Collective communication,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Collective communication,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.384160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.022282Z digest=sha256:a22a6737fc98e7b3820b735baa7e71157e18956d7bf4ac2382c6d9fe0e99f301

Observation 0a933052-8573-4252-86f6-55e8dd9aedf4 · outbound

This paper cites Towards the ideal on-chip fabric for 1-to-many and many-to-1 communication,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Towards the ideal on-chip fabric for 1-to-many and many-to-1 communication,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:32.102965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.127464Z digest=sha256:f32b844f5e370b62a1d3743bcad421fafaefb44a2448c9679df2612f81be7aeb

Observation decf9542-b5a8-411a-8e29-25f720ce5dcc · outbound

This paper cites GVSoC: a highly configurable, fast and accurate full- platform simulator for RISC-V based IoT processors,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators GVSoC: a highly configurable, fast and accurate full- platform simulator for RISC-V based IoT processors,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:31.742199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.250138Z digest=sha256:15b84f7cb151afddd08a83cc6f9d58b8c9d2c86764d43a42bc6d469a9e0277fc

Observation 8d73a607-c1f1-4753-be67-0d6e73a8d4f1 · outbound

This paper cites Snitch: A tiny pseudo dual-issue processor for area and energy efficient execution of floating-point intensive workloads,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Snitch: A tiny pseudo dual-issue processor for area and energy efficient execution of floating-point intensive workloads,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:31.429126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.357031Z digest=sha256:fd01cc08f79f8d3f58a9f20564f41a18b799039e6e296df576b4d186a82853c1

Observation 7c482cb8-b5c9-422f-a249-418a79543af7 · outbound

This paper cites Spatz: Clustering compact RISC-V-based vector units to maximize computing efficiency,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Spatz: Clustering compact RISC-V-based vector units to maximize computing efficiency,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:31.190544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.485665Z digest=sha256:9f3daaa2d730bb70714a4cdc1f193e39c87c6e7195e253c1389d3b773701ef6c

Observation c91be4fd-2578-4fca-9f12-9b05c8bff576 · outbound

This paper cites A high-performance, energy-efficient modular DMA engine architecture,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators A high-performance, energy-efficient modular DMA engine architecture,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:30.874902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.604764Z digest=sha256:e4a9d762d822a1a8a9961769af8e645fc37323713ce2221edaa3110eddf27a9f

Observation 1c013230-326d-4b0a-b056-7555a1d7a0ba · outbound

This paper cites RedMule: A mixed-precision matrix–matrix oper- ation engine for flexible and energy-efficient on-chip linear algebra and TinyML training acceleration,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators RedMule: A mixed-precision matrix–matrix oper- ation engine for flexible and energy-efficient on-chip linear algebra and TinyML training acceleration,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:30.533129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.720003Z digest=sha256:a4fbd92d7d6faa2c3f5ca6efb43cb83ccf91db739f320b21c3f172e778c449b0

Observation c09561ed-bd44-40a2-9f50-dbc1fa446dcd · outbound

This paper cites FlooNoC: A 645-Gb/s/link 0.15-pJ/B/hop open-source NoC with wide physical links and end-to-end AXI4 parallel multistream support,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlooNoC: A 645-Gb/s/link 0.15-pJ/B/hop open-source NoC with wide physical links and end-to-end AXI4 parallel multistream support,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:30.278817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.830392Z digest=sha256:9ccb63fbabeed751764f0d38642724f2032b8c47a2f491fb0ce2973f6fabc541

Observation f01a241a-6578-4f5a-a103-07e2a786ee47 · outbound

This paper cites DRAMSys: a flexible DRAM subsystem design space exploration framework,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators DRAMSys: a flexible DRAM subsystem design space exploration framework,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:29.966934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:28.952136Z digest=sha256:93a4804e190241866fb482c5b1670002df4d98b946e9ebe5f1f235b9c390744a

Observation d0180082-2dda-450a-a015-df5fd2fa1707 · outbound

This paper cites SUMMA: Scalable universal matrix multiplication algorithm,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators SUMMA: Scalable universal matrix multiplication algorithm,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:29.651742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:29.047456Z digest=sha256:47b7e8db5ef704a33c92309d540f487152b8472b570687897076b39c021264fe

Observation 6d21fc12-e934-4bfc-924e-d552d33a1e7e · outbound

This paper cites MI300X vs H100 vs H200 Benchmark Part 1: Training,.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators MI300X vs H100 vs H200 Benchmark Part 1: Training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:29.425767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:29.161352Z digest=sha256:e632377951623542d97f7e7b719e7faa52f8bd22ba8a719cf14e533ec441aa57

Pith citing papers

No inbound Pith citation observations are available.