Arctic Inference introduces Shift Parallelism, dynamic switching between tensor and sequence parallelism, achieving faster LLM inference and higher embedding throughput in a single deployment.
Ulysses: Unlocking low- latency, high-throughput inference for long context LLMs,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
extension 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
extension 1polarities
extend 1representative citing papers
citing papers explorer
-
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
Arctic Inference introduces Shift Parallelism, dynamic switching between tensor and sequence parallelism, achieving faster LLM inference and higher embedding throughput in a single deployment.