Pith. sign in

REVIEW 2 cited by

MNN: A Universal and Efficient Inference Engine

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.12418 v1 pith:QMD5NODC submitted 2020-02-27 cs.CV cs.DCcs.LG

classification cs.CVcs.DCcs.LG
keywords engineefficientinferencemobilechallengesdeepdeviceslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deploying deep learning models on mobile devices draws more and more attention recently. However, designing an efficient inference engine on devices is under the great challenges of model compatibility, device diversity, and resource limitation. To deal with these challenges, we propose Mobile Neural Network (MNN), a universal and efficient inference engine tailored to mobile applications. In this paper, the contributions of MNN include: (1) presenting a mechanism called pre-inference that manages to conduct runtime optimization; (2)deliveringthorough kernel optimization on operators to achieve optimal computation performance; (3) introducing backend abstraction module which enables hybrid scheduling and keeps the engine lightweight. Extensive benchmark experiments demonstrate that MNN performs favorably against other popular lightweight deep learning frameworks. MNN is available to public at: https://github.com/alibaba/MNN.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference

    cs.AR 2026-07 conditional novelty 7.0 of 10

    Cross-layer measurements of five mobile LLM frameworks on CPU/GPU/NPU reveal amplified NPU framework gaps, a prefill–decode backend phase split, and up to ~55% NPU energy savings from scheduling fixes.

  2. MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices

    cs.LG 2025-06 conditional novelty 4.0 of 10

    MNN-LLM, a mobile LLM inference engine based on MNN, reports up to 8.6x faster prefill than llama.cpp on a smartphone CPU through quantization, hybrid DRAM-Flash storage, and hardware-tuned kernels.

Pith tools