Serving frontier MoE and multimodal models on Ascend 910 via vLLM-Ascend is feasible but dominated by engineering cost from incomplete operators, fragile parallelism, kernel faults, and weak observability.
Gomez, Łukasz Kaiser, and Illia Polosukhin
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
Refined DHS targets, two-stage image-quality screening, and spherical-harmonic geo-encoding reduce KidSat MAE from 0.2167 to 0.1759 (18.83 percent relative) and reach 0.1658 on 33 African countries.
FastOCR dynamically selects a small subset of visual tokens per decoding step using focal-guided pruning and cross-step reuse, retaining 98% accuracy on Qwen2.5-VL while attending to only 5% of tokens and cutting attention latency by 3x.
LLM-generated, validity-filtered reward programs selected only on sparse Overcooked returns improve MAPPO coordination over the sparse baseline, especially under handoff and congestion bottlenecks.
citing papers explorer
-
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend
Serving frontier MoE and multimodal models on Ascend 910 via vLLM-Ascend is feasible but dominated by engineering cost from incomplete operators, fragile parallelism, kernel faults, and weak observability.
-
Enhancing the KidSat Model: Integrating Geographical Encoding and Data Quality Assessment for Childhood Poverty Prediction
Refined DHS targets, two-stage image-quality screening, and spherical-harmonic geo-encoding reduce KidSat MAE from 0.2167 to 0.1759 (18.83 percent relative) and reach 0.1658 on 33 African countries.
-
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
FastOCR dynamically selects a small subset of visual tokens per decoding step using focal-guided pruning and cross-step reuse, retaining 98% accuracy on Qwen2.5-VL while attending to only 5% of tokens and cutting attention latency by 3x.
-
Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning
LLM-generated, validity-filtered reward programs selected only on sparse Overcooked returns improve MAPPO coordination over the sparse baseline, especially under handoff and congestion bottlenecks.