Pith. sign in

REVIEW 4 major objections 5 minor 91 references

HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference

T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read HeteroMosaic argues that edge LLM inference on unified-memory chips should be scheduled as dependency-preserving micro-batches whose device assignment and timing are co-optimized around the measured critical path, yielding up to about 2x la

desk verdict Solid systems paper with a plausible central claim—heterogeneity-first scheduling beats single-accelerator baselines on edge SoCs—but the headline speedups need statistical backing and the HeteroInfer comparison is a re-creation, not the original. read the letter →

arxiv 2607.12839 v4 pith:776P64MB submitted 2026-07-14 cs.DC cs.AR

classification cs.DCcs.AR
keywords heterogeneousschedulingedgeLLMinferencemicro-batchingcausalattentionrooflinemodelunifiedmemorycritical-pathoptimizationenergyefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that on modern edge chips that pack a CPU, an integrated GPU, and a neural processor behind one memory system, the winning move is not to pick one accelerator and place every operation on it, but to treat inference as a scheduling problem: split the prompt into causally ordered chunks, run chunks that no longer depend on each other on different accelerators at once, and then tune each chunk's device assignment and timing around the measured critical path. It claims this heterogeneity-first approach, applied to unchanged off-the-shelf models, improves end-to-end time-to-first-token by up to roughly 1.7x to 2x over strong single-accelerator runtimes and cuts energy by up to about 45% on balanced systems. If true, edge devices could serve longer prompts and larger models within the same power envelope without changing model weights or quantization. The paper also shows that decode benefits are smaller because token generation is memory-bound, so the main opportunity is in prefill.

What carries the argument

The central machinery is causal parallel micro-batching: splitting prefill into sequential chunks where each chunk writes keys and values into the growing key-value cache and later chunks attend only through that cache, which is exact because attention is causal and which exposes overlap absent from the monolithic graph. A heterogeneous roofline model—a min-over-accelerators, max-over-splits expression with efficiency-scaled compute and bandwidth—bounds when that overlap is worth pursuing. Around the measured critical path, a latency-shaping policy edits each node's split ratio, accelerator assignment, or inserted bubble, accepting edits only when they improve re-measured end-to-end latency;

What would settle it

On the same SoC, take a fixed model and prompt, force every micro-batch onto the strongest single accelerator under the same DVFS and thermal state, and compare time-to-first-token, total energy, and generated-token equality against HeteroMosaic's heterogeneous schedule; if the single-accelerator run wins on latency or energy, or if the token streams diverge, the central claim fails.

Watch

Extended reading notes

Core claim

The paper's central discovery is that heterogeneous execution becomes a practical winning strategy for edge LLM prefill when it is exposed and scheduled, not merely assigned. The key move is causal parallel micro-batching: a long prompt is split into sequential chunks, each chunk extends the key-value cache, and later chunks attend only through that cache, which makes the decomposition mathematically exact for causal attention. This restructured graph exposes overlap that is invisible in a monolithic prefill, such as the post-attention work of one chunk running concurrently with the attention of the next chunk on a different accelerator. A trace-guided critical-interval co-optimization then

Load-bearing premise

The load-bearing premise is that splitting a long prompt into sequential chunks and letting later chunks attend to earlier chunks through the growing key-value cache leaves the model's output numerically unchanged—if chunked attention drifts from monolithic prefill, the reported speedups come with unstated changes in model behavior.

Editorial extensions

If this is right

  • If the central claim holds, prefill time-to-first-token is the main win: up to 1.73x over a strong iGPU-only path, 1.78x over an NPU-only path, and 2.05x over a portable open-source runtime on a balanced SoC, all with unchanged models and standard weight-only quantization.
  • Energy per inference falls by up to 45.3% because HeteroMosaic completes more useful work inside roughly the same peak-power envelope on TDP-constrained chips.
  • Decode tokens-per-second improves only modestly (up to about 13%) and only when the iGPU does not already saturate memory bandwidth; on stronger iGPUs heterogeneous decode is neutral or slightly harmful.
  • The scheduling principle is intended to generalize beyond LLMs to other transformer-based pipelines, and the trace-guided co-optimization runs offline once per model/device, so deployment cost is a configuration lookup rather than per-request search.
  • Because micro-batching also shrinks peak memory, the method extends to very long prompts, such as 70B-parameter models at 16K tokens, where monolithic prefill fails with out-of-memory errors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editor's inference: the reported 45% energy saving is measured as total wall-clock energy at a fixed power cap, so it partly reflects finishing sooner; battery-life gains would additionally depend on idle power draw during the saved time, which the paper does not model.
  • Editor's inference: a decisive testable extension is to compare logits or decoded tokens from monolithic prefill versus micro-batched prefill at 16K context; if numerical drift is negligible, the exactness claim is stronger than typical float-reordering behavior would suggest.
  • Editor's inference: the same micro-batching-plus-critical-path recipe may transfer to multi-stage transformer workloads such as vision-language or diffusion models, where modality-specific stages create additional independent work; the paper names this direction but leaves it open.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes HeteroMosaic, a heterogeneity-first scheduler for edge LLM inference on AMD Ryzen AI SoCs. It contributes (i) a Gables-style heterogeneous roofline model for identifying when iGPU+NPU execution can beat single-device bounds; (ii) a causal micro-batching formulation that decomposes prefill into dependency-preserving chunks to expose cross-accelerator overlap; (iii) a trace-guided critical-interval co-optimizer that jointly tunes micro-batch schedule, tensor-split allocation, and bubbles under DVFS/NPU queueing effects; and (iv) an AMD-specific runtime with shared-memory buffers, HIP/IRON kernels, and NPU queue coalescing. The evaluation covers three Ryzen AI SKUs, four AWQ W4A16 models, prompt lengths from 1K to 16K, and compares against iGPU/NPU baselines, llama.cpp, a HeteroInfer-style baseline, and production frameworks. The central claims are up to 1.73x/1.78x/2.05x speedups on the balanced SoC and up to 45.3% energy reduction.

Significance. If the quantitative claims hold, this is a substantive contribution: it separates heterogeneity-first scheduling from model restructuring/sparsity/quantization changes, which are confounded in several prior edge-LLM systems, and it demonstrates consistent end-to-end gains across three balance points. Strengths include a code artifact, broad platform/model coverage, 1024-run microbenchmarks with error bars, component ablations, and an honest treatment of decode headroom. The roofline model is not circular: its efficiency factors are calibrated from separate microbenchmarks and prior GEMM studies, then compared with, not fit to, end-to-end results. However, the end-to-end experiments that carry the headline numbers currently lack the statistical support the authors themselves show is necessary, and the comparison against "prior heterogeneous edge AI solutions" rests on a reconstructed baseline with an artificial iGPU throttle. The central idea is defensible, but the empirical evidence needs strengthening before the specific performance claims can be accepted.

major comments (4)
  1. [Sec. 6.1, 6.2; Fig. 8; Table 2] The paper's own GEMM microbenchmarks (Sec. 6.1) are averaged over 1024 runs and show error bars that "primarily reflect SoC power-management behavior" and turbo-to-sustained transitions. Yet the end-to-end prefill plots in Fig. 8, which carry the abstract's headline speedups, show only point estimates with no repeat count, error bars, or run-to-run distribution; the decode table (Table 2) also reports no variance. Because Algorithm 2 (Sec. 5.2.4) tunes configurations by executing them on the same model/platform/prompt ranges used in evaluation, the reported gains may reflect in-sample selection or favorable turbo draws. Please report number of repetitions, variance/confidence intervals for all end-to-end latency and energy numbers, and use a separate validation phase (e.g., held-out prompt lengths or a separate tuning run) to demonstrate that tuned schedules generalize.
  2. [Sec. 6.2, Footnote 4; Abstract] The HeteroInfer comparison is a re-creation that forces a 1:10 iGPU:NPU imbalance by "inserting bubbles on the iGPU path." This is not a faithful port of HeteroInfer's scheduler to AMD hardware; it is an iGPU-throttled tensor-partitioning baseline. Consequently, the abstract's claim of up to 2.35x speedup over "prior heterogeneous edge AI solutions" conflates heterogeneous-scheduling gains with the effect of artificially slowing the iGPU. The evaluation should either (a) compare against an unmodified tensor-partitioning baseline at the native AMD balance point, or (b) relabel the baseline as "HeteroInfer-like under an emulated 1:10 imbalance" and avoid presenting the speedup as a head-to-head improvement over HeteroInfer.
  3. [Sec. 6.3; Fig. 9; Abstract] The 45.3% energy reduction is a headline claim, but the energy methodology is underreported. Figure 9 shows 10 Hz current traces and labels "45.3% less J", yet the text does not state how energy was integrated from current, how many repeated runs were performed, what the variance was, or how thermal/turbo state was controlled across runs. Given the paper's own emphasis on turbo-to-sustained DVFS transitions, energy should be reported with the same statistical rigor as the latency microbenchmarks. This is load-bearing because the energy claim appears in the abstract and in the tokens-per-watt framing.
  4. [Sec. 5.2.4; Algorithm 2] Algorithm 2 accepts a proposed edit only if it improves measured end-to-end latency in the same evaluation setting (lines 13-16), with a search budget B and about eight hours of offline tuning per model/device configuration. This is an empirical minimization over a large configuration space. The paper should report the number of evaluated schedules/edits (B), show tuning curves or convergence, and evaluate the final configuration on settings not used during tuning. Without such validation, the comparison against fixed heuristic baselines is not fully controlled, and the reported "up to" numbers may reflect the result of exhaustive search rather than a generalizable scheduling policy.
minor comments (5)
  1. [Sec. 4, Eq. (2)-(5)] The relationship between per-accelerator bandwidth B_i and the shared B_peak is not fully specified. It should be clarified whether B_i are per-device bandwidth limits to the unified memory or are already derived from B_peak; otherwise the double-counting concern in the min expression is hard to evaluate. The CPU is present in the equations but omitted from Fig. 4; justify or quantify the omission.
  2. [Fig. 8] With five overlapping curves, the subplots are hard to read, especially in black-and-white. Consider panel-specific markers or separating the llama.cpp curve into a distinct style, and add repeat counts or error bars as discussed in the major comments.
  3. [Sec. 6.5, Fig. 11] The cross-framework comparison mixes numerical formats (integer, BFP16, BF16) and is appropriately labeled as coarse. It would help to state explicitly that Fig. 11 is not used to support the abstract's speedup claims, or to remove it from the narrative of the headline results.
  4. [Sec. 5.2.4] The eight-hour trace-guided search cost is mentioned but no details are given for the hardware, the search budget B, or the representative prompt-length ranges. A short reproducibility note with these values would make the offline tuning step more useful to follow-on work.
  5. [Sec. 4] The NPU efficiency ceiling of 0.4 is taken from references [81-83], some of which share authors with this paper. A sensitivity sweep over this value would make the roofline projections more robust and avoid the appearance of calibration by fiat.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the roofline is an analytical bound, end-to-end results are measured after offline search, and the only self-citation (NPU efficiency from ref [82]) is non-load-bearing.

full rationale

The paper's central claim is empirical rather than a derivation that reduces to its own inputs. The heterogeneous roofline model (Sec. 4) takes peak TOPS, bandwidth, operational intensity, and efficiency factors (eta_i, alpha_i) as inputs and computes an analytical upper bound; it is explicitly framed as a 'target' rather than an end-to-end predictor, and the paper states that measured microbenchmarks and traces later calibrate how far real execution falls below this ceiling. The end-to-end speedups in Fig. 8 are measured after an offline trace-guided search (Algorithm 2) that accepts edits only when re-executed latency improves; this is configuration search on the evaluated workloads, not a fitted parameter renamed as a prediction. The causal micro-batching exactness claim (Sec. 2.3) is a mathematical property of causal attention with an incrementally extended KV cache, not a definition of the target result. The only self-citation of note is the NPU efficiency 0.4 in Sec. 4, attributed to refs [81-83], of which ref [82] (Wang et al.) shares authors (Pang, Jun, Chen) with this paper; however, this value only sets a conservative upper bound for an illustrative speedup surface and does not enter the measured end-to-end claims, so it is a minor self-citation that is not load-bearing. Statistical robustness of the headline numbers (no error bars in Fig. 8 despite DVFS/turbo variation noted in Sec. 6.1 microbenchmarks) is a correctness/evidence concern, not a circularity one.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests primarily on empirical tuning (trace-guided search) and platform-specific runtime engineering. The analytical roofline is motivating and uses efficiency parameters imported from prior literature. No new physical entities are posited.

free parameters (6)
  • eta_i (compute efficiency of NPU/iGPU) = NPU ~0.4, iGPU ~0.7 (from refs [20,81-83])
    Used to scale published peak TOPS to effective compute in the heterogeneous roofline model (Section 4). Values are assumed/imported, not derived from the target result.
  • alpha_i (memory bandwidth efficiency) = not numerically specified
    Introduced in the roofline model to account for unified-memory contention, cache behavior, and runtime overheads (Section 4); no concrete values or measurement procedure are given.
  • micro-batch schedule (sizes and asymmetry) = searched offline per model/prompt-range (Algorithm 2)
    Micro-batch sizes are a decision variable in the search; the final schedule is fitted to measured latency on the evaluated workloads.
  • tensor-parallel split ratios and balance points = e.g., turbo 0.5, sustained 0.6 (Algorithm 1 comments)
    Split ratios and target balance points are selected empirically from microbenchmarks and tuned via trace-guided search; they are platform/model-specific fitted choices.
  • number of work-dispatch threads = 2 for AI 7/9, 3 for Max+ 395
    Hand-picked tuning parameter for CPU dispatch parallelism (Section 5.2.1).
  • search budget B = unspecified
    Algorithm 2's outer loop limit B controls how much trace-guided tuning is performed; not reported.
assumptions (5)
  • domain assumption Causal micro-batching preserves the exact mathematical result of monolithic prefill.
    Section 2.3 states the decomposition is exact because later micro-batches attend through the KV cache; this is load-bearing for the correctness of the schedule, but numerical equivalence is not formally proven.
  • standard math Gables-style roofline formulation models concurrent accelerators on a shared memory system.
    Section 4 builds on Gables [37] with per-device attainability min(B_i*I_i, P_i)/f_i; this is an accepted analytical model.
  • domain assumption NPU datapath reconfiguration is expensive and DMA reconfiguration is cheap.
    Section 5.3.2 cites prior work [75] and the paper's own measurements; the scheduling relies on coalescing to avoid datapath reconfiguration.
  • domain assumption Unified memory allows both accelerators to access shared weights without duplication or explicit copies.
    Section 5.3.1 describes DRM-based shared buffers; the entire weight-sharing benefit depends on this platform property.
  • standard math Attention scales as O(N^2 d) and projections as O(N d^2).
    Section 5.2.3 uses these complexity scaling laws to justify asymmetric micro-batching for long contexts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference." pith.science (2026). https://pith.science/paper/776P64MB

@misc{pith2026260712839,
  author       = {Pith},
  title        = {Pith review of: HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/776P64MB}},
  note         = {Machine review of arXiv:2607.12839}
}
read the original abstract

Modern edge system-on-chips (SoCs) combine CPUs, integrated GPUs (iGPUs), and neural processing units (NPUs), yet existing LLM runtimes typically make coarse device-level decisions or optimize operators in isolation. As a result, they underutilize heterogeneous resources, particularly on unified-memory platforms where performance depends on both device placement and task-graph coordination. We present HeteroMosaic, a heterogeneity-first scheduling framework for edge LLM inference. HeteroMosaic first uses a heterogeneous roofline model to identify when combining iGPU and NPU execution is beneficial. It then decomposes inference into dependency-preserving micro-batches that expose cross-accelerator overlap and applies trace-guided co-optimization of scheduling and device allocation under practical effects such as memory contention, DVFS, device variation, and NPU runtime overheads. We implement HeteroMosaic in PyTorch C++ and evaluate it on three AMD Ryzen AI platforms spanning NPU-heavy, balanced, and iGPU-heavy designs. On the balanced platform, HeteroMosaic achieves up to 1.73X speedup over an iGPU baseline, 1.78X over an NPU baseline, and 2.05X over frameworks such as llama dot cpp, while reducing energy by up to 45.3%. It also improves performance over prior heterogeneous edge AI solutions by up to 2.35X.

Figures

Figures reproduced from arXiv: 2607.12839 by the authors.

Figure 1
Figure 1. Overview of HeteroMosaic. Using off-the-shelf LLMs, HeteroMosaic restructures execution into dependency-preserving [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Different tensor partitioning schemes in GEMM. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of the heterogeneous roofline model. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Projected heterogeneous speedup over the iGPU [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Illustration of Heterogeneous Micro-Batching. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Example of Asymmetric Micro-Batching Schedules. [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Speedup of heterogeneous execution for GEMMs of varying sizes on the NPU and iGPU, normalized to the iGPU [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: This setting is substantially more challenging than the [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 8
Figure 8. Figure 8: End-to-end LLM prompt latency speedup across AMD Ryzen [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Measured current (A) over time during the same [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 11
Figure 11. Figure 11: Cross-framework comparison for Llama3-8B [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

91 extracted references · 15 linked inside Pith

  1. [1]

    Technical Report

    2024.Matrix Multiplication Performance Guide. Technical Report. NVIDIA. https://docs.nvidia.com/deeplearning/performance/dl-performance- matrix-multiplication/index.html

  2. [2]

    2025.AMD Ryzen™AI 9 HX 370

    Advanced Micro Devices, Inc. 2025.AMD Ryzen™AI 9 HX 370. https://www.amd.com/en/products/processors/laptop/ryzen/ai-300- series/amd-ryzen-ai-9-hx-370.html

  3. [3]

    Advanced Micro Devices, Inc. 2026. AMD Zen Deep Neural Network Library. https://www.amd.com/en/developer/zendnn.html. Accessed: 2026-05-17

  4. [4]

    Gulavani, and Ramachandran Ramjee

    Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2024. Taming Throughput-Latency Trade- off in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 117–136

  5. [5]

    AMD. [n. d.]. ROCm. https://github.com/ROCm

  6. [6]

    2023.How to accelerate AI applications on RDNA 3 using WMMA

    AMD. 2023.How to accelerate AI applications on RDNA 3 using WMMA. https: //gpuopen.com/learn/wmma_on_rdna3/ Describes RDNA 3 WMMA instructions and supported data types (FP16, BF16, INT8, INT4). Accessed: 2025-11-02

  7. [7]

    AMD. 2025. AMD XDNA Architecture. https://www.amd.com/en/technologies/ xdna.html. Accessed: 2025-07-28

  8. [8]

    AMD. 2026. AMD Ryzen™AI 7 350. https://www.amd.com/en/products/ processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-7-350.html. Product page, accessed 2026-04-02

Show all 91 references
  1. [9]

    AMD. 2026. AMD Ryzen™AI Max+ 395. https://www.amd.com/en/products/ processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html. Prod- uct page, accessed 2026-04-02

  2. [10]

    AMD. 2026. OnnxRuntime GenAI (OGA) Flow — Ryzen AI Software. https: //ryzenai.docs.amd.com/en/latest/hybrid_oga.html. Accessed: 2026-03-23

  3. [11]

    2024.HIP Runtime API: hipEventRecord

    AMD ROCm Team. 2024.HIP Runtime API: hipEventRecord. https://rocm.docs.amd.com/projects/HIP/en/develop/doxygen/html/group___ event.html#ga5df2309c9f29ca4c8e669db658d411b4 Accessed: 2025-07-28

  4. [12]

    2024.HIP Runtime API: hipStreamWaitValue32

    AMD ROCm Team. 2024.HIP Runtime API: hipStreamWaitValue32. https://rocm.docs.amd.com/projects/HIP/en/docs-develop/reference/hip_ runtime_api/modules/stream_memory_operations.html Accessed: 2025-07-28

  5. [13]

    2025.HIP Runtime API: Streams and Synchro- nization

    AMD ROCm Team. 2025.HIP Runtime API: Streams and Synchro- nization. https://rocm.docs.amd.com/projects/HIP/en/latest/understand/ programming_model.html Accessed: 2025-07-28

  6. [14]

    Anthropic. 2025. Claude Code. https://www.anthropic.com. AI coding assistant built on Claude; accessed 31 Jul 2025

  7. [15]

    Apple. [n. d.]. The Most Powerful Neural Engine Ever. https://www.apple.com/ newsroom/2024/05/apple-introduces-m4-chip

  8. [16]

    Apple. 2026. Core ML | Apple Developer Documentation. https://developer.apple. com/documentation/coreml. Accessed: 2026-03-23

  9. [17]

    Apple Inc. 2026. Metal | Apple Developer Documentation. https://developer. apple.com/documentation/metal. Accessed: 2026-05-17

  10. [18]

    Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu

  11. [19]

    Black Forest Labs. 2024. FLUX. https://github.com/black-forest-labs/flux. Official inference repository for FLUX.1 models

  12. [20]

    Cade Brown, Ahmad Abdelfattah, Stanimire Tomov, and Jack Dongarra. 2020. Design, Optimization, and Benchmarking of Dense Linear Algebra Algorithms on AMD GPUs. In2020 IEEE High Performance Extreme Computing Conference (HPEC). 1–7. https://doi.org/10.1109/HPEC43674.2020.9286214

  13. [21]

    Le Chen, Dahu Feng, Erhu Feng, Yingrui Wang, Rong Zhao, Yubin Xia, Pinjie Xu, and Haibo Chen. 2025. Characterizing Mobile SoC for Accelerating Hetero- geneous LLM Inference. InProceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. https://doi.org/10.1145...

  14. [22]

    Compute Express Link Consortium. [n. d.]. Compute Express Link Specification. CXL Consortium Specification. https://www.computeexpresslink.org Accessed: 2026-06-17

  15. [23]

    Fu, Stefano Ermon, Atri Rudra, and Christopher R’e

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R’e. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. InAdvances in Neural Information Processing Systems, Vol. 35. 16344–16359. arXiv:2205.14135 [cs.LG] https://arxiv.org/abs/2...

  16. [24]

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M"uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. 2024. Scaling Rectified Flow...

  17. [25]

    FastFlowLM. 2026. FastFlowLM. https://github.com/FastFlowLM/FastFlowLM. GitHub repository, accessed 2026-04-04

  18. [26]

    FLAME Project. 2026. BLIS: BLAS-like Library Instantiation Software Framework. https://github.com/flame/blis. Accessed: 2026-05-17

  19. [27]

    GGML-ORG. [n. d.]. llama.cpp. https://github.com/ggml-org/llama.cpp

  20. [28]

    ggml-org. 2024. What’s the difference between batch-size and ubatch-size? Discussion #6328. https://github.com/ggml-org/llama.cpp/discussions/6328. Ac- cessed: 2026-03-24

  21. [29]

    ggml-org. 2026. ggml: Tensor library for machine learning. https://github.com/ ggml-org/ggml. Accessed: 2026-05-17

  22. [30]

    ggml-org. 2026. llama.cpp Build Documentation. https://github.com/ggml-org/ llama.cpp/blob/master/docs/build.md. Accessed: 2026-05-17

  23. [31]

    ggml-org. 2026. llama.cpp common.h. https://github.com/ggml-org/llama.cpp/ blob/master/common/common.h. Accessed: 2026-03-24

  24. [32]

    ggml-org. 2026. llama.cpp Quantization Tool Documentation. https://github. com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md. Accessed: 2026-05-17

  25. [33]

    ggml-org. 2026. llama.cpp RPC Backend Documentation. https://github.com/ ggml-org/llama.cpp/blob/master/tools/rpc/README.md. Accessed: 2026-05-17

  26. [34]

    glassescrab. 2026. mlir-aie: An Open-Source Fork of the IRON API and MLIR- Based AI Engine Toolchain. https://github.com/glassescrab/mlir-aie. GitHub repository, accessed 2026-03-27

  27. [35]

    Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang, Ting Cao, and Ju Ren. 2026. Scaling LLM Test-Time Compute with Mobile NPU on Smartphones. InProceedings of the 21st European Conference on Computer Systems (EuroSys ’26). 2157–2172. https://doi.org/...

  28. [36]

    Jyothi Hariharan, Rahul Rama Varior, and Sunil Karunakaran. 2023. Real-time Driver Monitoring Systems on Edge AI Device.arXiv preprint arXiv:2304.01555 (2023). https://doi.org/10.48550/arXiv.2304.01555

  29. [37]

    Mark Hill and Vijay Janapa Reddi. 2019. Gables: A roofline model for mobile socs. In2019 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 317–330

  30. [38]

    Huawei. 2026. CANN: Ascend AI Computing Platform. https://www.hiascend. com/en/cann. Accessed: 2026-05-17

  31. [39]

    Erika Hunhoff, Joseph Melber, Kristof Denolf, Andra Bisca, Samuel Bayliss, Stephen Neuendorffer, Jeff Fifield, Jack Lo, Pranathi Vasireddy, Phil James- Roxby, and Eric Keller. 2025. Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface. arXi...

  32. [40]

    Intel. [n. d.]. Quick overview of Intel’s Neural Processing Unit (NPU). https: //intel.github.io/intel-npu-acceleration-library/npu.html

  33. [41]

    Intel. 2026. OpenVINO Documentation. https://docs.openvino.ai/. Accessed: 2026-05-17

  34. [42]

    kernel.org. [n. d.]. DRM Memory Management. https://www.kernel.org/doc/ html/v4.15/gpu/drm-mm.html

  35. [43]

    Khronos Group. 2026. OpenCL Registry. https://registry.khronos.org/OpenCL/. Accessed: 2026-05-17

  36. [44]

    Khronos Group. 2026. SYCL 2020 Specification. https://registry.khronos.org/ SYCL/specs/sycl-2020/html/sycl-2020.html. Accessed: 2026-05-17

  37. [45]

    Khronos Group. 2026. Vulkan Specification. https://registry.khronos.org/vulkan/ specs/latest/html/vkspec.html. Accessed: 2026-05-17

  38. [46]

    Hyungyo Kim, Nachuan Wang, Qirong Xia, Jinghan Huang, Amir Yazdanbakhsh, and Nam Sung Kim. 2025. LIA: A Single-GPU LLM Inference Acceleration with Cooperative AMX-Enabled CPU-GPU Computation and CXL Offloading. In Proceedings of the 52nd Annual International Symposium on Compu...

  39. [47]

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakr- ishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2024....

  40. [48]

    Muyang Li, Yujun Lin, Zhekai Zhang, et al. 2024. SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models. https://arxiv.org/abs/ 2411.05007

  41. [49]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Chen Wang, Wei-Ming Chen, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2023. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. arXiv preprint arXiv:2306.00978(2023)

  42. [50]

    Meta. [n. d.]. Llama 3. https://ai.meta.com/blog/meta-llama-3

  43. [51]

    Netlib. 2026. BLAS: Basic Linear Algebra Subprograms. https://www.netlib.org/ blas/. Accessed: 2026-05-17

  44. [52]

    NVIDIA. 2025. NVIDIA Isaac GR00T N1. https://developer.nvidia.com/isaac/gr00t. Official product page

  45. [53]

    NVIDIA. 2026. CUDA Toolkit Documentation. https://docs.nvidia.com/cuda/. Accessed: 2026-05-17

  46. [54]

    NVIDIA. 2026. NVIDIA NVLink and NVLink Switch. NVIDIA Data Center Tech- nology Overview. https://www.nvidia.com/en-us/data-center/nvlink/ Accessed: 2026-06-17

  47. [55]

    NVIDIA, Johan Bjorck, Fernando Casta neda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic...

  48. [56]

    OpenAI. 2023. ChatGPT. https://openai.com/chatgpt. Large-language-model conversational agent; accessed 31 Jul 2025

  49. [57]

    OpenAI. 2023. GPT-4 Technical Report. (2023). arXiv:2303.08774 [cs.CL] https: //arxiv.org/abs/2303.08774

  50. [58]

    2023.Khan Academy

    OpenAI. 2023.Khan Academy. https://openai.com/index/khan-academy/ An- nounces GPT-4 powering Khanmigo as a tutor and classroom assistant

  51. [59]

    Muhammad Osama, Duane Merrill, Cris Cecka, Michael Garland, and John D. Owens. 2023. Stream-K: Work-Centric Parallel Decomposition for Dense Matrix- Matrix Multiplication on the GPU. InProceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel ...

  52. [60]

    OWON Technology. 2026. OWON SPE Series 1 CH 100W–300W DC Power Sup- ply. https://www.owon.com.hk/products_owon_spe_series_1_ch_100w-300w_ dc_power_supply. Product page, accessed 2026-04-03

  53. [61]

    PCI-SIG. 2025. PCI Express Base Specification Revision 7.0. PCI-SIG Specification Library. https://pcisig.com/specifications Accessed: 2026-06-17

  54. [62]

    William Peebles and Saining Xie. 2023. Scalable Diffusion Models with Transform- ers. InProceedings of the IEEE/CVF International Conference on Computer Vision. 4195–4205. https://doi.org/10.48550/arXiv.2212.09748 arXiv:2212.09748 [cs.CV]

  55. [63]

    Physical Intelligence. 2025. Open Sourcing 𝜋0. https://www.pi.website/blog/ openpi. Project blog post

  56. [64]

    Physical Intelligence. 2025. openpi: Open-Source Models and Packages for Ro- botics. https://github.com/Physical-Intelligence/openpi. GitHub repository

  57. [65]

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin Le...

  58. [66]

    Physical Intelligence, Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li- Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch...

  59. [67]

    PyTorch. 2025. PyTorch. https://pytorch.org/. Accessed: 2025-07-28

  60. [68]

    Qualcomm. [n. d.]. A new era of possibility with on-device AI. https://www. qualcomm.com/products/technology/artificial-intelligence

  61. [69]

    Qualcomm. 2026. Llama-v2-7B-Chat. https://aihub.qualcomm.com/models/ llama_v2_7b_chat?domain=Generative+AI&useCase=Text+Generation& chipsets=qualcomm-snapdragon-8gen3. Qualcomm AI Hub model card, accessed 2026-03-27

  62. [70]

    Qualcomm Technologies, Inc. 2023. Snapdragon 8 Gen 3 Mobile Plat- form. https://www.qualcomm.com/smartphones/products/8-series/snapdragon- 8-gen-3-mobile-platform. Accessed: 2026-05-17

  63. [71]

    Qualcomm Technologies, Inc. 2026. Qualcomm AI Engine Direct SDK Documen- tation. https://docs.qualcomm.com/nav/home/QNN_general_overview.html? product=1601111740009302. Accessed: 2026-03-23

  64. [72]

    Qualcomm Technologies, Inc. 2026. Qualcomm Hexagon SDK Documentation. https://docs.qualcomm.com/nav/home?product=1601111740010422. Accessed: 2026-05-17

  65. [73]

    Qualcomm Technologies, Inc. 2026. Quantization — Qualcomm AI Engine Direct SDK Documentation. https://docs.qualcomm.com/bundle/publicresource/topics/ 80-63442-10/quantization.html. Accessed: 2026-03-26

  66. [74]

    Qualcomm Technologies, Inc. 2026. Writing QNN HTP Op Package — Qualcomm AI Engine Direct SDK Documentation. https://docs.qualcomm.com/bundle/ publicresource/topics/80-63442-10/writing_op_package.html. Accessed: 2026- 03-26

  67. [75]

    André Rösti and Michael Franz. 2025. Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools. In33rd IEEE Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) 2025, Fayetteville, AR, USA, May 4–...

  68. [76]

    Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. FlexGen: high-throughput generative inference of large language models with a single GPU. InProceedings of the 40th International Confere...

  69. [77]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.arXiv preprint arXiv:1909.08053(2019). https://doi.org/10.48550/arXiv.1909.08053 arXi...

  70. [78]

    Yushan Siriwardhana, Pawani Porambage, Madhusanka Liyanage, and Mika Ylianttila. 2021. A Survey on Mobile Augmented Reality With 5G Mobile Edge Computing: Architectures, Applications, and Technical Aspects.IEEE Commu- nications Surveys & Tutorials23, 2 (2021), 1160–1192. https...

  71. [79]

    Zhiyi Song et al . 2018. Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design. InAAAI Conference on Artificial Intelligence

  72. [80]

    Nazish Tahir and Ramviyas Parasuraman. 2025. Edge Computing and its Application in Robotics: A Survey.arXiv preprint arXiv:2507.00523(2025). https://doi.org/10.48550/arXiv.2507.00523

  73. [81]

    Endri Taka, Andre Roesti, Joseph Melber, Pranathi Vasireddy, Kristof Denolf, and Diana Marculescu. 2025. Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs. arXiv:2512.13282 [cs.AR] https://arxiv. org/abs/2512.13282 19 Jun et al. HeteroMosa...

  74. [82]

    Chengyue Wang, Wesley Pang, Xinrui Wu, Gregory Jun, Luis Romero, Endri Taka, Diana Marculescu, Tony Nowatzki, Pranathi Vasireddy, Joseph Melber, Deming Chen, and Jason Cong. 2025. Can Asymmetric Tile Buffering Be Beneficial? arXiv:2511.16041 [cs.DC] https://arxiv.org/abs/2511.16041

  75. [83]

    Erwei Wang, Samuel Bayliss, Andra Bisca, Zachary Blair, Sangeeta Chowdhary, Kristof Denolf, Jeff Fifield, Brandon Freiberger, Erika Hunhoff, Phil James-Roxby, Jack Lo, Joseph Melber, Stephen Neuendorffer, Eddie Richter, André Rosti, Javier Setoain, Gagandeep Singh, Endri Taka,...

  76. [84]

    Xubin Wang, Zhiqing Tang, Jianxiong Guo, Tianhui Meng, Chenhao Wang, Tian Wang, and Weijia Jia. 2025. Empowering Edge Intelligence: A Comprehensive Survey on On-Device AI Models.Comput. Surveys57, 9 (2025), 1–39. https: //doi.org/10.1145/3724420

  77. [85]

    Gonzalez

    Yichuan Wang, Zhifei Li, Shu Liu, Yongji Wu, Ziming Mao, Yilong Zhao, Xiao Yan, Zhiying Xu, Yang Zhou, Ion Stoica, Sewon Min, Matei Zaharia, and Joseph E. Gonzalez. 2025. LEANN: A Low-Storage Vector Index. https://doi.org/10.48550/ arXiv.2506.08276 arXiv:2506.08276 [cs.IR]

  78. [86]

    Wang and Duen Horng Chau

    Zijie J. Wang and Duen Horng Chau. 2024. MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. https://doi.org/10.1145/3626772.3657662

  79. [87]

    Patterson

    Samuel Williams, Andrew Waterman, and David A. Patterson. 2009. Roofline: An Insightful Visual Performance Model for Multicore Architectures.Commun. ACM52, 4 (April 2009), 65–76. https://doi.org/10.1145/1498765.1498785

  80. [88]

    World Wide Web Consortium. 2026. WebGPU Specification. https://www.w3. org/TR/webgpu/. Accessed: 2026-05-17

  81. [89]

    Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Gang Huang, Mengwei Xu, and Xuanzhe Liu. 2025. Fast On-Device LLM Inference with NPUs. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (ASPL...

  82. [90]

    Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, and Haibo Chen. 2024. PowerInfer-2: Fast Large Language Model Inference on a Smartphone.arXiv preprint arXiv:2406.06282(2024). https://doi.org/10.48550/arXiv.2406.06282 20

  83. [2023]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    All are Worth Words: A ViT Backbone for Diffusion Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22669– 22679. https://doi.org/10.48550/arXiv.2209.12152 arXiv:2209.12152 [cs.CV]

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.