Tevatron 3.0 adds Megatron expert parallelism to the open Tevatron reranker toolkit, trains a 30B MoE reranker on a small cluster, and shows it matches dense-8B reranking quality at higher serving throughput.
An analysis of the softmax cross entropy loss for learning-to-rank with binary relevance
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Tevatron Meets Megatron: Expert-Parallel LLM Reranker Training on an Academic Budget
Tevatron 3.0 adds Megatron expert parallelism to the open Tevatron reranker toolkit, trains a 30B MoE reranker on a small cluster, and shows it matches dense-8B reranking quality at higher serving throughput.