Pith. sign in

REVIEW 1 cited by

Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00025 v2 pith:JNDKID67 submitted 2024-01-05 cs.DC cs.AI

classification cs.DCcs.AI
keywords matrixfusedimprovementinferencekernelaveragedecompositionfound
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose an implementation of an efficient fused matrix multiplication kernel for W4A16 quantized inference, where we perform dequantization and GEMM in a fused kernel using a SplitK work decomposition. Our implementation shows improvement for the type of skinny matrix-matrix multiplications found in foundation model inference workloads. In particular, this paper surveys the type of matrix multiplication between a skinny activation matrix and a square weight matrix. Our results show an average of 65% speed improvement on A100, and an average of 124% speed improvement on H100 (with a peak of 295%) for a range of matrix dimensions including those found in a llama-style model, where m < n = k.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference

    cs.AR 2026-08 conditional novelty 6.0 of 10

    Dual-sparse LLM decoding can be accelerated by an RLC-CSC spMspV kernel, and a small SIMT-core hardware addition is proposed to remove the remaining index-reconstruction and accumulation bottlenecks.

Pith tools