Pith. sign in

REVIEW 11 cited by

MobileNetV4 -- Universal Models for the Mobile Ecosystem

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.10518 v2 pith:XHZXEQVU submitted 2024-04-16 cs.CV

classification cs.CV
keywords mobilemnv4modelssearchacceleratorsaccuracyarchitectureblock
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the latest generation of MobileNets, known as MobileNetV4 (MNv4), featuring universally efficient architecture designs for mobile devices. At its core, we introduce the Universal Inverted Bottleneck (UIB) search block, a unified and flexible structure that merges Inverted Bottleneck (IB), ConvNext, Feed Forward Network (FFN), and a novel Extra Depthwise (ExtraDW) variant. Alongside UIB, we present Mobile MQA, an attention block tailored for mobile accelerators, delivering a significant 39% speedup. An optimized neural architecture search (NAS) recipe is also introduced which improves MNv4 search effectiveness. The integration of UIB, Mobile MQA and the refined NAS recipe results in a new suite of MNv4 models that are mostly Pareto optimal across mobile CPUs, DSPs, GPUs, as well as specialized accelerators like Apple Neural Engine and Google Pixel EdgeTPU - a characteristic not found in any other models tested. Finally, to further boost accuracy, we introduce a novel distillation technique. Enhanced by this technique, our MNv4-Hybrid-Large model delivers 87% ImageNet-1K accuracy, with a Pixel 8 EdgeTPU runtime of just 3.8ms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SnapGen is a 379M-parameter UNet with cross-architecture distillation and a 1.38M-parameter decoder that generates 1024x1024 images on a phone in about 1.4 seconds, with GenEval 0.66 and ImageNet FID 2.06.

  2. EMOv2: Pushing 5M Vision Model Frontier

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 5M-parameter backbone with shared-weight spanning window attention sets new accuracy records across classification, detection, and generation benchmarks.

  3. Efficient Track Anything

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A lightweight video segmentation model with a vanilla ViT encoder and pooled memory cross-attention matches SAM 2 closely while running twice as fast and using 2.4x fewer parameters.

  4. FDIO: Frequency Decomposed Inertial Odometry

    cs.CV 2025-11 conditional novelty 5.0 of 10

    Splitting pedestrian IMU signals into smooth and jumpy frequency bands — Mamba on the smooth band, multi-scale convolutions on the jumpy band — cuts average trajectory error by roughly a third versus the RoNIN ResNet ...

  5. Toroidal area-preserving parameterizations of genus-one closed surfaces

    math.NA 2025-08 unverdicted novelty 5.0 of 10

    Four Riemannian optimization algorithms (projected/Riemannian gradient and conjugate gradient) are proposed to compute toroidal area-preserving parameterizations by minimizing stretch energy on a power manifold of ring tori.

  6. MoNetV2: Enhanced Motion Network for Freehand 3D Ultrasound Reconstruction

    eess.IV 2025-06 conditional novelty 5.0 of 10

    MoNetV2 improves freehand 3D ultrasound reconstruction by adding multi-level consistency losses and a multi-modal self-supervised strategy that reduce cumulative drift.

  7. iFormer: Integrating ConvNet and Transformer for Mobile Application

    cs.CV 2025-01 conditional novelty 5.0 of 10

    iFormer combines a mobile-tuned ConvNeXt backbone with single-head modulation attention, reaching 80.4% ImageNet top-1 accuracy at 1.10 ms iPhone 13 latency.

  8. PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A lightweight refiner with a coarse-to-fine denoising module, noise-based pretraining, and a scale-shift invariant gradient-matching loss achieves state-of-the-art high-resolution metric depth with up to 10x faster inference.

  9. CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Adding nearest-neighbor and cross nearest-neighbor supervision from frozen pretrained unimodal encoders to the CLIP loss improves lightweight vision-language models on zero-shot and retrieval benchmarks.

  10. OCDet: Object Center Detection via Bounding Box-Aware Heatmap Prediction on Edge Devices with NPUs

    cs.CV 2024-11 conditional novelty 5.0 of 10

    OCDet predicts object center heatmaps with Generalized Centerness and Balanced Continuous Focal Loss, and reports higher recall and CAS than YOLO11 on edge NPUs.

  11. STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification

    cs.CV 2025-09 conditional novelty 4.0 of 10

    STA-Net, a 401K-parameter model with a decoupled shape-texture attention module, reaches 89.00% accuracy and 88.96% F1 on the CCMT plant disease dataset.

Pith tools