Pith. sign in

Rethinking llm inference bottlenecks: Insights from latent attention and mixture-of-experts.arXiv preprint arXiv:2507.15465, 2025

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

years

2026 4 2025 2

representative citing papers

Think Before You Grid-Search: Floor-First Triage for LLM Serving

cs.PF · 2026-07-07 · conditional · novelty 6.0 · 2 refs

LLM serving should triage by five-resource analytical floors and wall ordering, not grid search; on 16×H20, TP16 is capacity-capped at ~70 while EP+DP attention reaches ~644 concurrent 8K requests.

Janus: Disaggregating Attention and Experts for Scalable MoE Inference

cs.DC · 2025-12-15 · unverdicted · novelty 6.0

JANUS disaggregates attention and MoE layers onto separate GPU pools with an expert-balancing scheduler and SLO-aware scaling, delivering up to 4.7x higher per-GPU throughput than prior MoE systems under token-level latency constraints.

citing papers explorer

Showing 6 of 6 citing papers.