REVIEW 11 cited by
MobileNetV4 -- Universal Models for the Mobile Ecosystem
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present the latest generation of MobileNets, known as MobileNetV4 (MNv4), featuring universally efficient architecture designs for mobile devices. At its core, we introduce the Universal Inverted Bottleneck (UIB) search block, a unified and flexible structure that merges Inverted Bottleneck (IB), ConvNext, Feed Forward Network (FFN), and a novel Extra Depthwise (ExtraDW) variant. Alongside UIB, we present Mobile MQA, an attention block tailored for mobile accelerators, delivering a significant 39% speedup. An optimized neural architecture search (NAS) recipe is also introduced which improves MNv4 search effectiveness. The integration of UIB, Mobile MQA and the refined NAS recipe results in a new suite of MNv4 models that are mostly Pareto optimal across mobile CPUs, DSPs, GPUs, as well as specialized accelerators like Apple Neural Engine and Google Pixel EdgeTPU - a characteristic not found in any other models tested. Finally, to further boost accuracy, we introduce a novel distillation technique. Enhanced by this technique, our MNv4-Hybrid-Large model delivers 87% ImageNet-1K accuracy, with a Pixel 8 EdgeTPU runtime of just 3.8ms.
Forward citations
Cited by 11 Pith papers
-
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
SnapGen is a 379M-parameter UNet with cross-architecture distillation and a 1.38M-parameter decoder that generates 1024x1024 images on a phone in about 1.4 seconds, with GenEval 0.66 and ImageNet FID 2.06.
-
EMOv2: Pushing 5M Vision Model Frontier
A 5M-parameter backbone with shared-weight spanning window attention sets new accuracy records across classification, detection, and generation benchmarks.
-
Efficient Track Anything
A lightweight video segmentation model with a vanilla ViT encoder and pooled memory cross-attention matches SAM 2 closely while running twice as fast and using 2.4x fewer parameters.
-
FDIO: Frequency Decomposed Inertial Odometry
Splitting pedestrian IMU signals into smooth and jumpy frequency bands — Mamba on the smooth band, multi-scale convolutions on the jumpy band — cuts average trajectory error by roughly a third versus the RoNIN ResNet ...
-
Toroidal area-preserving parameterizations of genus-one closed surfaces
Four Riemannian optimization algorithms (projected/Riemannian gradient and conjugate gradient) are proposed to compute toroidal area-preserving parameterizations by minimizing stretch energy on a power manifold of ring tori.
-
MoNetV2: Enhanced Motion Network for Freehand 3D Ultrasound Reconstruction
MoNetV2 improves freehand 3D ultrasound reconstruction by adding multi-level consistency losses and a multi-modal self-supervised strategy that reduce cumulative drift.
-
iFormer: Integrating ConvNet and Transformer for Mobile Application
iFormer combines a mobile-tuned ConvNeXt backbone with single-head modulation attention, reaching 80.4% ImageNet top-1 accuracy at 1.10 ms iPhone 13 latency.
-
PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation
A lightweight refiner with a coarse-to-fine denoising module, noise-based pretraining, and a scale-shift invariant gradient-matching loss achieves state-of-the-art high-resolution metric depth with up to 10x faster inference.
-
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance
Adding nearest-neighbor and cross nearest-neighbor supervision from frozen pretrained unimodal encoders to the CLIP loss improves lightweight vision-language models on zero-shot and retrieval benchmarks.
-
OCDet: Object Center Detection via Bounding Box-Aware Heatmap Prediction on Edge Devices with NPUs
OCDet predicts object center heatmaps with Generalized Centerness and Balanced Continuous Focal Loss, and reports higher recall and CAS than YOLO11 on edge NPUs.
-
STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
STA-Net, a 401K-parameter model with a decoupled shape-texture attention module, reaches 89.00% accuracy and 88.96% F1 on the CCMT plant disease dataset.
Discussion (0). Continue with ORCID to comment.