On a single H100, the 20.9B-parameter MoE model GPT-OSS-20B shows roughly 32% higher decode throughput, 26% lower energy per 1,000 tokens, and 32% lower peak VRAM than dense Qwen3-32B at 2,048-token context, at the cost of higher time-to-first-token.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
On a single H100, the 20.9B-parameter MoE model GPT-OSS-20B shows roughly 32% higher decode throughput, 26% lower energy per 1,000 tokens, and 32% lower peak VRAM than dense Qwen3-32B at 2,048-token context, at the cost of higher time-to-first-token.