Vortex provides a programmable frontend and backend for sparse attention in LLM serving, delivering up to 3.46x throughput over full attention while preserving accuracy.
https://arxiv.org/ abs/2411.13259
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Rust sparse kernels match Eigen and PSBLAS performance for CSC formats but trail PETSc's blocked CSR optimizations.
citing papers explorer
-
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents
Vortex provides a programmable frontend and backend for sparse attention in LLM serving, delivering up to 3.46x throughput over full attention while preserving accuracy.
-
Evaluating Rust for Sparse Matrix Kernels in Scientific Computing
Rust sparse kernels match Eigen and PSBLAS performance for CSC formats but trail PETSc's blocked CSR optimizations.