Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:29.161352Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.18824.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:29.161352Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6baf57a8-36f6-4226-92a5-69aa9eaf98e4 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators On the computational complexity of self-attention,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2da04287-b6d6-4094-9ce7-7ba8329aa926 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0cfb9b3-f0e9-4eff-95a5-e7f8e75b694e · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Data movement is all you need: A case study on optimizing transformers,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b9234b8b-47e6-4fe6-9478-77b618a41694 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention: Fast and memory-efficient exact attention with IO-awareness,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b1303b3-74ce-4003-a6e8-97045b74dd87 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention-2: Faster attention with better parallelism and work partitioning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff115aa5-8c1a-4601-a0ef-4a94f2508275 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead9f0f6-6a3c-4d01-8439-511de6f2d9fe · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators SambaNova SN40L: Scaling the AI memory wall with dataflow and composition of experts,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 555d3de4-b705-4e0b-9c89-cd24aeb2e39d · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Wafer-scale AI: GPU impossible performance,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ea76046-4229-47bc-9ca7-5f44509bda34 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Blackhole & TT-Metalium: The standalone AI computer and its programming model,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f5bedae7-44b5-429c-85de-efbe08a69ee3 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators 16.2 rngd: A 5nm tensor-contraction processor for power-efficient inference on large language models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47282922-b877-4fad-93a0-853e5b5f66bb · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Attention in SRAM on Tenstorrent Grayskull
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08011de-6c7d-4a4d-834e-bc71041f9341 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FLAT: An optimized dataflow for mitigating attention bottlenecks,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2bff695c-0266-4cd1-ad94-120f6114f8ee · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Fusemax: Leveraging extended einsums to optimize attention accelerator design,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb749651-27e1-4ec3-ba8e-74bbc933b1ae · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Gemini: Mapping and architecture co-exploration for large- scale DNN chiplet accelerators,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b3bf5695-480a-476c-99a9-b3a58ff851e5 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators DOJO: The microarchitecture of Tesla’s exa-scale computer,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05faa99a-a993-4ae1-af69-9b1534aa2b64 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Collective communication,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a933052-8573-4252-86f6-55e8dd9aedf4 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Towards the ideal on-chip fabric for 1-to-many and many-to-1 communication,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation decf9542-b5a8-411a-8e29-25f720ce5dcc · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators GVSoC: a highly configurable, fast and accurate full- platform simulator for RISC-V based IoT processors,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d73a607-c1f1-4753-be67-0d6e73a8d4f1 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Snitch: A tiny pseudo dual-issue processor for area and energy efficient execution of floating-point intensive workloads,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c482cb8-b5c9-422f-a249-418a79543af7 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators Spatz: Clustering compact RISC-V-based vector units to maximize computing efficiency,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c91be4fd-2578-4fca-9f12-9b05c8bff576 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators A high-performance, energy-efficient modular DMA engine architecture,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1c013230-326d-4b0a-b056-7555a1d7a0ba · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators RedMule: A mixed-precision matrix–matrix oper- ation engine for flexible and energy-efficient on-chip linear algebra and TinyML training acceleration,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c09561ed-bd44-40a2-9f50-dbc1fa446dcd · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators FlooNoC: A 645-Gb/s/link 0.15-pJ/B/hop open-source NoC with wide physical links and end-to-end AXI4 parallel multistream support,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f01a241a-6578-4f5a-a103-07e2a786ee47 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators DRAMSys: a flexible DRAM subsystem design space exploration framework,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d0180082-2dda-450a-a015-df5fd2fa1707 · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators SUMMA: Scalable universal matrix multiplication algorithm,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d21fc12-e934-4bfc-924e-d552d33a1e7e · outbound
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators MI300X vs H100 vs H200 Benchmark Part 1: Training,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.