Pith. sign in

Paper Citation Record · LEDGER

How to keep pushing ML accelerator performance? Know your rooflines!

As of 7 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2505.16346.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16346 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:30.520235Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy60
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c44066fd-fa82-459e-ac9a-e92d9805bfd9 · outbound

This paper cites Visualizing size of large language models,.

How to keep pushing ML accelerator performance? Know your rooflines! Visualizing size of large language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.398472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.398472Z digest=sha256:b5243c1fe7649447e416fe2b3903bd2508a9869f7687ed82375fe07e8e037bc8

Observation 14d25844-d249-4f4c-977e-e25c5c79aa5c · outbound

This paper cites Trends in deep learning hardware,.

How to keep pushing ML accelerator performance? Know your rooflines! Trends in deep learning hardware,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:32.138137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.403476Z digest=sha256:900ebc01ba9563f9ac5fe2e8dd053ec9bb5bf2feb22d5875ac02cefd364c208d

Observation 7086185c-8b13-4bc9-9546-2fd5afc2a679 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures,.

How to keep pushing ML accelerator performance? Know your rooflines! Roofline: an insightful visual performance model for multicore architectures,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:32.113461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.408467Z digest=sha256:c42e20c05fc1fb06f54febdb82968267549a67e77b298c16e4b83284c667c780

Observation 83c26e29-7cd7-4049-b5c9-4c43790a0f65 · outbound

This paper cites Eyeriss: A apatial architecture for energy-efficient dataflow for convolutional neural networks,.

How to keep pushing ML accelerator performance? Know your rooflines! Eyeriss: A apatial architecture for energy-efficient dataflow for convolutional neural networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:32.096760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.412848Z digest=sha256:c40266c13efdc62f862e21ed080a9ed695a77c32367bb81a098ca5da13917a82

Observation dba73b76-278c-45fe-8064-49a7104e97cc · outbound

This paper cites Roofline performance analysis of dnn architectures on cpu and gpu systems,.

How to keep pushing ML accelerator performance? Know your rooflines! Roofline performance analysis of dnn architectures on cpu and gpu systems,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:32.079775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.419412Z digest=sha256:37a042d2448eb6cb7d67a5839273541eb23f3dd029e2c83bd4f10a20c2aa9133

Observation 3bc34efd-f60a-415e-b5a8-0faad797ffe0 · outbound

This paper cites an unresolved cited work.

How to keep pushing ML accelerator performance? Know your rooflines! Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:06:32.048841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.424820Z digest=sha256:a740afe7c0bbe20a25a07608f8b076ef6db6959c01614fc108f24110252efca3

Observation 5ebbde3f-f00f-4f35-a464-fe03ea9f0e38 · outbound

This paper cites Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,.

How to keep pushing ML accelerator performance? Know your rooflines! Understanding reuse, performance, and hardware cost of dnn dataflow: A data-centric approach,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:32.028318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.438721Z digest=sha256:f868a5434e7ba99e627cbef4383eccae2ad8831db59b9e6217d36cbd0fb38b6f

Observation 191aa578-ee30-4289-a75c-2836b4ec5a45 · outbound

This paper cites A roofline model of energy,.

How to keep pushing ML accelerator performance? Know your rooflines! A roofline model of energy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:32.008663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.443914Z digest=sha256:c0a463981503a439f35c11a010495f0ca8b8de2beac705f79a4f93cede12bcdf

Observation 2db46215-a0bb-4c48-84f7-8a8d5ad0049d · outbound

This paper cites Symphony: Orchestrating sparse and dense tensors with hierarchical heterogeneous processing,.

How to keep pushing ML accelerator performance? Know your rooflines! Symphony: Orchestrating sparse and dense tensors with hierarchical heterogeneous processing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.991855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.449801Z digest=sha256:1c48fa6ef098b0cd83d7a96a1a7bdfa7d5772668d904ae136cc70962a58fb71b

Observation 79867109-ebbf-4d1d-8b49-b0c9761378ec · outbound

This paper cites Lots of questions on Google’s “Trillium.

How to keep pushing ML accelerator performance? Know your rooflines! Lots of questions on Google’s “Trillium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.976054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.455217Z digest=sha256:15e995ede53cd5f562bc9423ed63c2e7916b7f90edd81c63b4cae179f14bd68b

Observation fcbb83b6-24a5-426b-b612-a27c077aa0f7 · outbound

This paper cites Envi- sion: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm fdsoi,.

How to keep pushing ML accelerator performance? Know your rooflines! Envi- sion: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm fdsoi,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.959854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.461094Z digest=sha256:4dbbef82ab846ff6f4310a8e6f0abc985f4888943cf6b9e9509a4278c5c045c5

Observation 03919dc5-ba5d-4c84-a46c-846a3f0ba0cc · outbound

This paper cites 9.5 a 6k-mac feature-map-sparsity-aware neural processing unit in 5nm flagship mobile soc,.

How to keep pushing ML accelerator performance? Know your rooflines! 9.5 a 6k-mac feature-map-sparsity-aware neural processing unit in 5nm flagship mobile soc,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.939247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.468026Z digest=sha256:484264064999290437f39504b0ecfcb2111a1ae16ff056e192159c0cd8ba97f9

Observation 563c3ace-a04a-40e0-b028-5fc467f91985 · outbound

This paper cites Compute solution for tesla’s full self-driving computer,.

How to keep pushing ML accelerator performance? Know your rooflines! Compute solution for tesla’s full self-driving computer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.921995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.472900Z digest=sha256:1898246bca847e5c5fdd950ec6c9157b9be96e2a8c7a1e1579a3e794a52499b4

Observation a7c94c0d-2efb-4268-8c84-19ebc1307c38 · outbound

This paper cites 7.2 a 12nm programmable convolution-efficient neural- processing-unit chip achieving 825tops,.

How to keep pushing ML accelerator performance? Know your rooflines! 7.2 a 12nm programmable convolution-efficient neural- processing-unit chip achieving 825tops,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.902241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.477473Z digest=sha256:e8e801a2233f41cbfd0ca1964a982c1d357532b4dd6ac86e67d2c69c7e31943a

Observation d74fe782-4893-4413-82b4-4950a2811a22 · outbound

This paper cites Groq rocks neural networks,.

How to keep pushing ML accelerator performance? Know your rooflines! Groq rocks neural networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.883652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.484255Z digest=sha256:005319067f08c6672e02e31838fe511fe4423deec804781abdf273a1df39b81d

Observation db943f5f-3a0c-4ed3-b6d0-bdf3d7e97261 · outbound

This paper cites 9.1 a 7nm 4-core ai chip with 25.6tflops hybrid fp8 training, 102.4tops int4 inference and workload-aware throttling,.

How to keep pushing ML accelerator performance? Know your rooflines! 9.1 a 7nm 4-core ai chip with 25.6tflops hybrid fp8 training, 102.4tops int4 inference and workload-aware throttling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.859940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.491187Z digest=sha256:082e062d4bbb5936afd4d1d86c26764cdf03e7297aaf3267a0ca82d22fde3877

Observation 21688fde-3472-4885-9133-5e67b0f4bbc7 · outbound

This paper cites 16.7 a 40-310tops/w sram-based all-digital up to 4b in-memory computing multi-tiled nn accelerator in fd-soi 18nm for deep-learning edge applications,.

How to keep pushing ML accelerator performance? Know your rooflines! 16.7 a 40-310tops/w sram-based all-digital up to 4b in-memory computing multi-tiled nn accelerator in fd-soi 18nm for deep-learning edge applications,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.841139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.501017Z digest=sha256:cb12eea79a0ea5c5323713dc10967b03d62e2e4940bd95e594874699e5ea3467

Observation 64e534e9-7d1c-4f72-aba6-e2a83c65919d · outbound

This paper cites Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architecture,.

How to keep pushing ML accelerator performance? Know your rooflines! Charm: Composing heterogeneous accelerators for matrix multiply on versal acap architecture,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.512744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.512744Z digest=sha256:8b8b979ec1904e6b0adc4c0eb891639ab253727cd7aab94622d65e583e76d0c3

Observation 074e189b-9dd1-4e3d-bd31-b2e6bb4362eb · outbound

This paper cites Davinci: A scalable architecture for neural network computing,.

How to keep pushing ML accelerator performance? Know your rooflines! Davinci: A scalable architecture for neural network computing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.821177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.523587Z digest=sha256:7663d8a740d4f4c9a0a6807df61e5ed908a1b508da4b9b4369f905793289adb8

Observation 70b5df40-95fa-47b8-b99e-193403ae7e95 · outbound

This paper cites Nvidia tensor core programmability, performance & precision,.

How to keep pushing ML accelerator performance? Know your rooflines! Nvidia tensor core programmability, performance & precision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.538744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.538744Z digest=sha256:a28fedad666193d562701cde6be1d2fb256c03d860723a2bf7dfd2d9995d8fc6

Observation a7876e8c-e647-4b9b-b903-617d5ecc35ae · outbound

This paper cites A charge domain sram compute-in-memory macro with c-2c ladder- based 8-bit mac unit in 22-nm finfet process for edge inference,.

How to keep pushing ML accelerator performance? Know your rooflines! A charge domain sram compute-in-memory macro with c-2c ladder- based 8-bit mac unit in 22-nm finfet process for edge inference,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.557111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.557111Z digest=sha256:4ddc62ca82ab61b57ce926329e8ab04596e48f6465928fa607386e1e044e678f

Observation 8aa39835-d1be-4fdd-968a-736471b5bba6 · outbound

This paper cites A 22 nm, 1540 top/s/w, 12.1 top/s/mm 2 in-memory analog matrix-vector-multiplier for dnn acceleration,.

How to keep pushing ML accelerator performance? Know your rooflines! A 22 nm, 1540 top/s/w, 12.1 top/s/mm 2 in-memory analog matrix-vector-multiplier for dnn acceleration,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.775674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.584583Z digest=sha256:096171cc0f8b4e794fc5f52c1a07528a2bcb7a6537b71c6fc105065a12b16e2b

Observation 060dfa3f-6df2-41e0-8d36-efaafad155e2 · outbound

This paper cites A 64-tile 2.4- mb in-memory-computing cnn accelerator employing charge-domain compute,.

How to keep pushing ML accelerator performance? Know your rooflines! A 64-tile 2.4- mb in-memory-computing cnn accelerator employing charge-domain compute,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.752647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.613354Z digest=sha256:bec6461d2fe82c6c45ae184a0474d1f81115bd5c7e9b3c84465a2a0fd4972f82

Observation 0b0afbc5-c8bf-4721-b4fb-82278323b147 · outbound

This paper cites Compute Solution for Tesla’s Full Self-Driving Computer,.

How to keep pushing ML accelerator performance? Know your rooflines! Compute Solution for Tesla’s Full Self-Driving Computer,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.733368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.643554Z digest=sha256:f6ee121dbec836cfdf0a285fa8b3922df704bb6d8d8e61cb67602eaebaa1b102

Observation fae8ddfc-2b2d-4e3a-9f0c-497ee270e85d · outbound

This paper cites Hardware for deep learning,.

How to keep pushing ML accelerator performance? Know your rooflines! Hardware for deep learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.715995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.669402Z digest=sha256:1bf2dfc8cbc39ec38f4a37838db9ff4469216f88f3b8e94c09a93624aee6f895

Observation da16470b-2d3e-4fe0-84a9-b071dd26d719 · outbound

This paper cites Lincoln ai computing survey (laics) update,.

How to keep pushing ML accelerator performance? Know your rooflines! Lincoln ai computing survey (laics) update,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.701147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.691612Z digest=sha256:8d40f2f76a459ff105d22fef97164ccb0d67853c591d82a4e491559a3b369afb

Observation 901acd9a-d4f1-4ea0-ac01-b5ca945a06b6 · outbound

This paper cites Neural network accelerator comparison.

How to keep pushing ML accelerator performance? Know your rooflines! Neural network accelerator comparison

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.682603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.702506Z digest=sha256:d33e96343a94247a7ee96b9386ede536405e51d695b4da264b7a3baee363e919

Observation 97d55eaf-e1ea-4233-aa9f-aea58d0b400a · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

How to keep pushing ML accelerator performance? Know your rooflines! LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.707975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.707975Z digest=sha256:cf25bfe854104362146631f1bbd4b40d915fc712e3ae8b5241363f540ce39f70

Observation 08353473-84cd-4cb4-bd49-537fa186db6a · outbound

This paper cites Minifloats on risc-v cores: Isa extensions with mixed- precision short dot products,.

How to keep pushing ML accelerator performance? Know your rooflines! Minifloats on risc-v cores: Isa extensions with mixed- precision short dot products,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.666883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.712718Z digest=sha256:646fff6b1c96367f1d05a8fdb2309a691f99a0f8e345e477b8c166c284e36847

Observation 8555f8ba-09e2-4b4a-af27-409f6d015b63 · outbound

This paper cites Cutie: Beyond petaop/s/w ternary dnn inference acceleration with better-than-binary energy efficiency,.

How to keep pushing ML accelerator performance? Know your rooflines! Cutie: Beyond petaop/s/w ternary dnn inference acceleration with better-than-binary energy efficiency,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.651276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.719228Z digest=sha256:12eeb0ded4e6e93c810d0ff7c4fa7e40ae878820431291647601aec3a8157922

Observation 0e670c07-2879-40bf-b013-1f9a3a083047 · outbound

This paper cites Binareye: An always-on energy-accuracy-scalable binary cnn processor with all memory on chip in 28nm cmos,.

How to keep pushing ML accelerator performance? Know your rooflines! Binareye: An always-on energy-accuracy-scalable binary cnn processor with all memory on chip in 28nm cmos,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.635016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.727130Z digest=sha256:6c8747019bc4c788129450c003e5ce25394cb10a1fb8670c385d3a1bb33b5a6b

Observation 0a3287c9-08c5-4d5d-a841-4c814fd7ab59 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

How to keep pushing ML accelerator performance? Know your rooflines! BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.752659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.752659Z digest=sha256:4070195e8169b178e7593e93b5052b7d3d8c3fbfebd9f65ee09c117ea7956532

Observation 1519180e-97b3-4911-bfc5-9e5d87d183b3 · outbound

This paper cites A 3 tops/w risc-v parallel cluster for inference of fine-grain mixed-precision quantized neural networks,.

How to keep pushing ML accelerator performance? Know your rooflines! A 3 tops/w risc-v parallel cluster for inference of fine-grain mixed-precision quantized neural networks,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.615689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:29.907854Z digest=sha256:0056e942ea54c5c72a415913010e973c24d37629f47faf9b52cfef69dba5a5f8

Observation 0ba24f90-6cda-4cec-98bc-13dd8c9fa289 · outbound

This paper cites Marsellus: A heterogeneous risc-v ai-iot end-node soc with 2–8 b dnn acceleration and 30%-boost adaptive body biasing,.

How to keep pushing ML accelerator performance? Know your rooflines! Marsellus: A heterogeneous risc-v ai-iot end-node soc with 2–8 b dnn acceleration and 30%-boost adaptive body biasing,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.601169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.045702Z digest=sha256:4a1d178691e23637250e92e790020b76588093ceb585fa6c798db2fa6abb88f4

Observation 99a00f12-8207-439c-b6f2-350163c31afb · outbound

This paper cites Microscaling Data Formats for Deep Learning.

How to keep pushing ML accelerator performance? Know your rooflines! Microscaling Data Formats for Deep Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.167879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.167879Z digest=sha256:6ee24331c094d0758a3d157d29e3352459494e2ecfbd41eb0e117dec96ebda57

Observation 2cf05c9b-d6e0-4f88-98ba-393fd718c14d · outbound

This paper cites Nvidia blackwell platform: Advancing generative ai and accelerated computing,.

How to keep pushing ML accelerator performance? Know your rooflines! Nvidia blackwell platform: Advancing generative ai and accelerated computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.585901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.260534Z digest=sha256:864ff0e42c4d5ce031ee9fbc38fb02d460df8a2575f23c1196d588f61e6ed430

Observation 1680c9ec-81fd-4344-8309-0f0dbff0fc78 · outbound

This paper cites Siracusa: A 16 nm heterogenous risc-v soc for extended reality with at-mram neural engine,.

How to keep pushing ML accelerator performance? Know your rooflines! Siracusa: A 16 nm heterogenous risc-v soc for extended reality with at-mram neural engine,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.571402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.295340Z digest=sha256:e8e6fb3fa8d682d9f4398d6e36151fd7bbc242a50e2613bc24c31b2b64860301

Observation 3fe2bff3-6cba-4d4f-9df5-377dcb967997 · outbound

This paper cites Onyx: A 12nm 756 gops/w coarse-grained reconfigurable array for accelerating dense and sparse applications,.

How to keep pushing ML accelerator performance? Know your rooflines! Onyx: A 12nm 756 gops/w coarse-grained reconfigurable array for accelerating dense and sparse applications,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.557095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.311228Z digest=sha256:3e511d16440d6155c7a1bfdabc8f56e97772c1e017ecd63523fae38aa836b981

Observation 811baf5b-cbc8-40a8-b6b8-58e74ad81a01 · outbound

This paper cites Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch.

How to keep pushing ML accelerator performance? Know your rooflines! Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.316475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.316475Z digest=sha256:35ad01ea3d489c5205258c31af2a21c2d4e9a1664c2c8fdfe8c13b830461c540

Observation fd4078a9-32f3-41e9-9fcb-d49e353faaa0 · outbound

This paper cites 3.2 the a100 datacenter gpu and ampere architecture,.

How to keep pushing ML accelerator performance? Know your rooflines! 3.2 the a100 datacenter gpu and ampere architecture,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.543279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.321479Z digest=sha256:76ac6c307989fda3db23d28dfb605a912f72e287185af3dcf9df30676a034792

Observation 1b9ff632-a182-4763-8a18-bfe3f114406f · outbound

This paper cites Venom: A vectorized n: M format for unleashing the power of sparse tensor cores,.

How to keep pushing ML accelerator performance? Know your rooflines! Venom: A vectorized n: M format for unleashing the power of sparse tensor cores,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.527977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.327263Z digest=sha256:1b8e19566983c089f836b073a8e291968aa9c0d5058dba8bfd0bd19cd87999f3

Observation 442b8fd4-7f49-4dfd-bff1-240ef09e13e4 · outbound

This paper cites Occamy: A 432-core dual-chiplet dual-hbm2e 768-dp-gflop/s risc-v system for 8- to-64-bit dense and sparse computing in 12-nm finfet,.

How to keep pushing ML accelerator performance? Know your rooflines! Occamy: A 432-core dual-chiplet dual-hbm2e 768-dp-gflop/s risc-v system for 8- to-64-bit dense and sparse computing in 12-nm finfet,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.511828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.332379Z digest=sha256:c8671cf1bbc7a6bd2c14542eaa8bd9859785afc97eb5ce962add248743432518

Observation a916b1ac-aac4-455c-b223-d85dd214cb28 · outbound

This paper cites Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,.

How to keep pushing ML accelerator performance? Know your rooflines! Neupims: Npu-pim heterogeneous acceleration for batched llm inferencing,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.338151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.338151Z digest=sha256:0bf70702fc6e9c8b473f59435d047495c87541e03a998c8991ab4ea68efaf906

Observation 17058cdc-295b-4200-86ac-6371834add2f · outbound

This paper cites Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures.

How to keep pushing ML accelerator performance? Know your rooflines! Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:06:30.773144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.344978Z digest=sha256:b1754be9aa0f24ad5444ad151b490e368e14c8e967492fe144e3176abfe70dcc

Observation c504234d-02c1-4508-a794-fdaf31d7348c · outbound

This paper cites In-memory computation of a machine-learning classifier in a standard 6t sram array,.

How to keep pushing ML accelerator performance? Know your rooflines! In-memory computation of a machine-learning classifier in a standard 6t sram array,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.496919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.350099Z digest=sha256:70da9652895e9cb264266f7ff291707e72d57e44cb4d3d1d946270f4136dd7bd

Observation eabae999-f60c-4fe5-bc52-73e6730bdc6a · outbound

This paper cites An energy-efficient memory-based high-throughput vlsi architecture for convolutional networks,.

How to keep pushing ML accelerator performance? Know your rooflines! An energy-efficient memory-based high-throughput vlsi architecture for convolutional networks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.479682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.354542Z digest=sha256:7b4eb5b27893eeb864d3c0e9d27fbe50da71184de544ec32b690fd42c8d17bbf

Observation dc024c9b-6911-468e-a603-4c7ed3dba6b1 · outbound

This paper cites Fast, energy-efficient, robust, and reproducible mixed-signal neuromorphic classifier based on embedded nor flash memory technology,.

How to keep pushing ML accelerator performance? Know your rooflines! Fast, energy-efficient, robust, and reproducible mixed-signal neuromorphic classifier based on embedded nor flash memory technology,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.457417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.359436Z digest=sha256:c15fc297e5ac05f888b341d94ef92cdf45c7d8266d153eec36d877a435cdd129

Observation 1983e91c-efc9-4b7e-9918-6b05794fdb74 · outbound

This paper cites Analog in-memory subthreshold deep neural network accelerator,.

How to keep pushing ML accelerator performance? Know your rooflines! Analog in-memory subthreshold deep neural network accelerator,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.441396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.364825Z digest=sha256:12f0515e666b7dba64023846e118548a6026e1999690f0209ac00b93d3787940

Observation a97abc10-7d55-49d8-8883-98a46308033d · outbound

This paper cites A 5-nm 254-tops/w 221-tops/mm2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,.

How to keep pushing ML accelerator performance? Know your rooflines! A 5-nm 254-tops/w 221-tops/mm2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage- frequency scaling and simultaneous mac and write operations,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.425267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.369284Z digest=sha256:f2cb2304afb2dd65a281b32cd0448a3ed7eb448008d1afb6266129f87b38be49

Observation dc0712c1-5f9f-4572-8f47-94f1c6f54a44 · outbound

This paper cites 16.4 an 89tops/w and 16.3tops/mm2 all-digital sram-based full-precision compute-in memory macro in 22nm for machine-learning edge applications,.

How to keep pushing ML accelerator performance? Know your rooflines! 16.4 an 89tops/w and 16.3tops/mm2 all-digital sram-based full-precision compute-in memory macro in 22nm for machine-learning edge applications,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.410366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.373554Z digest=sha256:e4b7391e5f7b1cb688581484757a69271fa92bd7e3b0196b0a400993309a1734

Observation db9547c9-63ff-48aa-9301-808349548048 · outbound

This paper cites A maximally row- parallel mram in-memory-computing macro addressing readout circuit sensitivity and area,.

How to keep pushing ML accelerator performance? Know your rooflines! A maximally row- parallel mram in-memory-computing macro addressing readout circuit sensitivity and area,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.395141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.378539Z digest=sha256:f9bc0efa06560eafa1539d4af14c83f1dc1e6740ce89ab5456e7d59e33b5cd60

Observation fd9d4dde-1f9e-452b-994a-62a1a717dc71 · outbound

This paper cites A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,.

How to keep pushing ML accelerator performance? Know your rooflines! A programmable heterogeneous microprocessor based on bit-scalable in-memory comput- ing,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.383254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.383254Z digest=sha256:3cc5b87f2589f8d21703755f988e0a5b1b98bab952d8c791f3f512b31bd57372

Observation 6393fdbe-e573-4a73-bd6a-fc229cb2005d · outbound

This paper cites A crossbar array of magnetoresistive memory devices for in-memory computing,.

How to keep pushing ML accelerator performance? Know your rooflines! A crossbar array of magnetoresistive memory devices for in-memory computing,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.387652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.387652Z digest=sha256:a5cba8062bb464ef994ed067a81567de0afab00b17424e6ec3fa7e08d3acf402

Observation 7cd66dde-bc47-49df-a0eb-f1c9e5eba12e · outbound

This paper cites In-memory computing: Advances and prospects,.

How to keep pushing ML accelerator performance? Know your rooflines! In-memory computing: Advances and prospects,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.392346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.392346Z digest=sha256:d3121871b6cff99b4ead1f00290ac0e98946725104c249826cb11ed6ac89519e

Observation 37ddb158-edd6-4020-8d8a-5accaacf7207 · outbound

This paper cites 14.2 a compute sram with bit-serial integer/floating-point operations for programmable in-memory vector acceleration,.

How to keep pushing ML accelerator performance? Know your rooflines! 14.2 a compute sram with bit-serial integer/floating-point operations for programmable in-memory vector acceleration,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.354015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.397269Z digest=sha256:88b6c8bc060b8a07e3fc4485d79d44b8bee72a85e706ed697049162d5dcaaf0c

Observation c4466db2-cdb4-4a65-b9c0-9ca6dfc39acd · outbound

This paper cites A 40nm 64kb 26.56tops/w 2.37mb/mm2rram binary/compute-in-memory macro with 4.23x im- provement in density and > 75% use of sensing dynamic range,.

How to keep pushing ML accelerator performance? Know your rooflines! A 40nm 64kb 26.56tops/w 2.37mb/mm2rram binary/compute-in-memory macro with 4.23x im- provement in density and > 75% use of sensing dynamic range,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.337684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.402092Z digest=sha256:6a1ddfa89f5be04d5b0e14a07f0d320745a1f2b16f1e8078c653122d697a27d8

Observation 04b325b6-8759-454e-9434-474b467266b1 · outbound

This paper cites Funda- mental limits on energy-delay-accuracy of in-memory architectures in inference applications,.

How to keep pushing ML accelerator performance? Know your rooflines! Funda- mental limits on energy-delay-accuracy of in-memory architectures in inference applications,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.318288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.406524Z digest=sha256:57df9e69475c11637746f741e133eebea5c9504cb1d8ac0f7f85312134500d4a

Observation 64a7278b-04c8-4782-a5a6-d54aab5e3df7 · outbound

This paper cites 11.3 metis aipu: A 12nm 15tops/w 209.6tops soc for cost- and energy-efficient inference at the edge,.

How to keep pushing ML accelerator performance? Know your rooflines! 11.3 metis aipu: A 12nm 15tops/w 209.6tops soc for cost- and energy-efficient inference at the edge,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.302813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.410836Z digest=sha256:acccd541ffa37b9926f619dacc798ff19d2ddb5cd7062c4f71cdd5916685bdfb

Observation 1f3e82c2-4fd0-45b6-be87-3f584cd7db5a · outbound

This paper cites Benchmarking in-memory computing architectures,.

How to keep pushing ML accelerator performance? Know your rooflines! Benchmarking in-memory computing architectures,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.415453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.415453Z digest=sha256:02ee7e5738147388b8944477524c89c6f6c4af5bb2f0386831f7100a12be73a1

Observation c2fb4c15-99b8-4560-b5be-7029727d1d2e · outbound

This paper cites A 22nm 128-kb mram row/column-parallel in-memory computing macro with memory- resistance boosting and multi-column adc readout,.

How to keep pushing ML accelerator performance? Know your rooflines! A 22nm 128-kb mram row/column-parallel in-memory computing macro with memory- resistance boosting and multi-column adc readout,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.278613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.420737Z digest=sha256:db96ab1c3be7d59e7aedb6a2f3068b66399889c137dc4236154c2f83eecdd93f

Observation 8af9691b-3084-4a53-9b91-dc7aaad9beaf · outbound

This paper cites A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference,.

How to keep pushing ML accelerator performance? Know your rooflines! A 64-core mixed-signal in-memory compute chip based on phase-change memory for deep neural network inference,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.425731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.425731Z digest=sha256:39efc68221197148cbd7cd21a8303ae82d89a53d675ff44c8277c07d8451959e

Observation e721fac0-ee09-4f4e-b3ae-d4b9c7e2f8d7 · outbound

This paper cites An n40 256k×44 embedded rram macro with sl-precharge sa and low-voltage current limiter to improve read and write performance,.

How to keep pushing ML accelerator performance? Know your rooflines! An n40 256k×44 embedded rram macro with sl-precharge sa and low-voltage current limiter to improve read and write performance,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.250233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.430911Z digest=sha256:3b6e771870e4a614c08ffe5ac9d14545080f1d600f545b79b5bd5274deaa9a6f

Observation 57ef3e98-bfbe-45af-9fe0-48b39b3e2f75 · outbound

This paper cites Cmos- embedded stt-mram arrays in 2x nm nodes for gp-mcu applications,.

How to keep pushing ML accelerator performance? Know your rooflines! Cmos- embedded stt-mram arrays in 2x nm nodes for gp-mcu applications,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.234950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.435960Z digest=sha256:95fc746fa49b6fa82f52af406b3f3fc9147e6f7ee62f4f1739dcace9781a14f6

Observation 847322d6-5d32-4d10-b9f9-6b0dceb45256 · outbound

This paper cites A switched-capacitor sram in-memory computing macro with high-precision, high-efficiency differential archi- tecture,.

How to keep pushing ML accelerator performance? Know your rooflines! A switched-capacitor sram in-memory computing macro with high-precision, high-efficiency differential archi- tecture,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.220998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.441328Z digest=sha256:c998f1e374f66ec7b0a0e7965c29c82011ca90eec7bd267fe9d3c72141d7ae78

Observation 499a3c28-1233-4957-b69f-4fa88b284e65 · outbound

This paper cites Scalable and Programmable Neural Network Inference Accelerator Based on In-Memory Computing,.

How to keep pushing ML accelerator performance? Know your rooflines! Scalable and Programmable Neural Network Inference Accelerator Based on In-Memory Computing,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.206563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.445701Z digest=sha256:a7ddbe1a8f60132c38f766c646aba8c11f0a27b011e1427c6a391d8ba11d7824

Observation a78fcd0f-2f3d-4bea-892a-66ee8ff32747 · outbound

This paper cites Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,.

How to keep pushing ML accelerator performance? Know your rooflines! Interstellar: Using Halide’s Scheduling Language to Analyze DNN Accelerators,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.450465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.450465Z digest=sha256:132a75eb30fce1e6c0fa6dcc1f95122237e64779dfd9c4f3cbc908776dc08e2e

Observation d5106f5f-7b45-4e8f-8825-d07f2c8fa3a4 · outbound

This paper cites MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,.

How to keep pushing ML accelerator performance? Know your rooflines! MAESTRO: A Data-Centric Approach to Understand Reuse, Performance, and Hardware Cost of DNN Mappings,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.192491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.454805Z digest=sha256:4f9232484aea520657352d7129e492438fc0cdaf018fb471873914bcda8b132e

Observation abdc799a-4fe5-46f0-9623-1de4d6bb730f · outbound

This paper cites Timeloop: A Systematic Approach to DNN Accelerator Evaluation,.

How to keep pushing ML accelerator performance? Know your rooflines! Timeloop: A Systematic Approach to DNN Accelerator Evaluation,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.179380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.459381Z digest=sha256:2464db259338fe2d1490a5e47a494ec34a9a44eddc258738f104ca965c863d48

Observation b8576801-923c-4513-acf6-f513d0ae90d5 · outbound

This paper cites ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,.

How to keep pushing ML accelerator performance? Know your rooflines! ZigZag: Enlarging Joint Architecture-Mapping Design Space Exploration for DNN Accelerators,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.163852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.463852Z digest=sha256:246cd4f8f9f1f7f6d35997e337a9f933ba1568501bedebb53874d454a9478b38

Observation 2cc2e601-e74a-4511-9e51-3ef07c37610c · outbound

This paper cites CoSA: Scheduling by constrained op- timization for spatial accelerators,.

How to keep pushing ML accelerator performance? Know your rooflines! CoSA: Scheduling by constrained op- timization for spatial accelerators,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.147252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.471518Z digest=sha256:440f460ac47ba5d7b7bf864947a9b52da06965fbb3396e2e7114433d9fcf3ff2

Observation b6ac7aae-c0da-4d9c-a770-63efc2c79417 · outbound

This paper cites Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search,.

How to keep pushing ML accelerator performance? Know your rooflines! Mind Mappings: Enabling Efficient Algorithm-Accelerator Mapping Space Search,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.476277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.476277Z digest=sha256:1d21d7981e32cfc6d9edd809fe54ca4036cef80c1745af172c91639f84a37a13

Observation 8ce0af98-da9c-489c-9b97-929ae3a283eb · outbound

This paper cites GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm,.

How to keep pushing ML accelerator performance? Know your rooflines! GAMMA: Automating the HW Mapping of DNN Models on Accelerators via Genetic Algorithm,

Reference 73

Resolution
verified exact
doi, observed 2026-08-07T15:06:30.566889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.480952Z digest=sha256:815ad0f39c68094de89fc86b6078ebe489a9721d53080079519bbdb11cc49d02

Observation 2c87e16f-c313-489e-91ac-527b88f5fe22 · outbound

This paper cites Stream: Design space exploration of layer-fused dnns on hetero- geneous dataflow accelerators,.

How to keep pushing ML accelerator performance? Know your rooflines! Stream: Design space exploration of layer-fused dnns on hetero- geneous dataflow accelerators,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.131537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.485112Z digest=sha256:06196b256597dfcc8a4adbed8ef3fcfc73bbb0a78d4ba8b5b99d59024300140e

Observation 8a588eea-947f-4fc9-90b8-ecc9ac56877d · outbound

This paper cites The groq software-defined scale-out tensor streaming multiprocessor : From chips-to-systems architectural overview,.

How to keep pushing ML accelerator performance? Know your rooflines! The groq software-defined scale-out tensor streaming multiprocessor : From chips-to-systems architectural overview,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.117852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.488902Z digest=sha256:39ca457d050f4029c4881a685d3c5c7cecee098f7aa62f91ad73daa1331834a6

Observation 17a54891-ec19-44ba-bd33-97e9311fa29c · outbound

This paper cites Application specific instruction processor based implementation of a gnss receiver on an fpga,.

How to keep pushing ML accelerator performance? Know your rooflines! Application specific instruction processor based implementation of a gnss receiver on an fpga,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.103230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.492972Z digest=sha256:4e4b1c31756c4dc14b1de645c0e4662682ddb82b27ea3adcefc96be0a98b39b5

Observation 1d680a15-8dd8-4544-a543-a003d2402db8 · outbound

This paper cites How flexible is your com- puting system?.

How to keep pushing ML accelerator performance? Know your rooflines! How flexible is your com- puting system?

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.088800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.497613Z digest=sha256:6c35f527590830f6af1f17a760fbea08eff48cc275e0e1586639049a0e20085e

Observation a9dedbe3-43ab-407c-aeec-1b8d2a4311e5 · outbound

This paper cites Tandem processor: Grappling with emerging operators in neural networks,.

How to keep pushing ML accelerator performance? Know your rooflines! Tandem processor: Grappling with emerging operators in neural networks,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.501738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.501738Z digest=sha256:77b470c7a84a37d569bd991ef66819347eea94319b47a9043115cf5656eec410

Observation cc7cdc0a-dd28-4d26-9de6-1a32daf336d9 · outbound

This paper cites Mec: memory-efficient convolution for deep neural network,.

How to keep pushing ML accelerator performance? Know your rooflines! Mec: memory-efficient convolution for deep neural network,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.063778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.505928Z digest=sha256:473b1acca6deaf0e881ae32598969299373c78ff7680d4cc61be340b9b19677e

Observation 196c9181-8b97-4077-97e0-4ae7c239699b · outbound

This paper cites A formalism of dnn accelerator flexibility,.

How to keep pushing ML accelerator performance? Know your rooflines! A formalism of dnn accelerator flexibility,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.048331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.510290Z digest=sha256:162d307db18ff32e3eb3c727b54dbeb59aecd925e7e16c8b07c41250735db24e

Observation 84c6dd27-560d-4be8-a390-6d1be9a0e729 · outbound

This paper cites Mlir: Scaling compiler infrastructure for domain specific computation,.

How to keep pushing ML accelerator performance? Know your rooflines! Mlir: Scaling compiler infrastructure for domain specific computation,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:30.515891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:30.515891Z digest=sha256:eef323212bf4589c2e46e67e732c3e30094be8c9f2b40ae87b28144dc2018a1b

Observation d484f9f9-f15a-4eb2-988f-a8ac376c7f42 · outbound

This paper cites The hardware lottery,.

How to keep pushing ML accelerator performance? Know your rooflines! The hardware lottery,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:31.023084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:30.520235Z digest=sha256:c5faa473fd58825932847291852663722c1c92d2069c6eae8ef468ca0776cdc9

Pith citing papers

No inbound Pith citation observations are available.