SpecGen introduces speculative generation to fork non-reasoning kernel candidates during LLM reasoning traces, enabling early termination and parallel profiling to reduce end-to-end optimization time on H200 GPUs.
SwizzlePerf: Hardware-aware LLMs for GPU kernel performance optimization
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
background 1representative citing papers
A diagnosis-driven evolutionary search with retrieval-augmented expert initialization improves Triton kernel correctness and speed on KernelBench, reaching 99–100% correctness on Level 2.
LEO is a cross-vendor GPU stall root-cause analyzer that traces stalled instructions backward through register and synchronization dependencies, achieving 1.73–1.82x geometric-mean speedups across 21 workloads on NVIDIA, AMD, and Intel GPUs.
citing papers explorer
-
SpecGen: Accelerating Agentic Kernel Optimization with Speculative Generation
SpecGen introduces speculative generation to fork non-reasoning kernel candidates during LLM reasoning traces, enabling early termination and parallel profiling to reduce end-to-end optimization time on H200 GPUs.
-
Kernel Foundry: A Diagnosis-driven Evolutionary Kernel Optimizer with Multi-Experts
A diagnosis-driven evolutionary search with retrieval-augmented expert initialization improves Triton kernel correctness and speed on KernelBench, reaching 99–100% correctness on Level 2.
-
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
LEO is a cross-vendor GPU stall root-cause analyzer that traces stalled instructions backward through register and synchronization dependencies, achieving 1.73–1.82x geometric-mean speedups across 21 workloads on NVIDIA, AMD, and Intel GPUs.