Pith. sign in

REVIEW 5 cited by

MobileViTv3: Mobile-Friendly Vision Transformer with Simple and Effective Fusion of Local, Global and Input Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.15159 v2 pith:FJUUVTYT submitted 2022-09-30 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords fusionmodelsblockcreateimagenet-1kade20kbetterdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

MobileViT (MobileViTv1) combines convolutional neural networks (CNNs) and vision transformers (ViTs) to create light-weight models for mobile vision tasks. Though the main MobileViTv1-block helps to achieve competitive state-of-the-art results, the fusion block inside MobileViTv1-block, creates scaling challenges and has a complex learning task. We propose changes to the fusion block that are simple and effective to create MobileViTv3-block, which addresses the scaling and simplifies the learning task. Our proposed MobileViTv3-block used to create MobileViTv3-XXS, XS and S models outperform MobileViTv1 on ImageNet-1k, ADE20K, COCO and PascalVOC2012 datasets. On ImageNet-1K, MobileViTv3-XXS and MobileViTv3-XS surpasses MobileViTv1-XXS and MobileViTv1-XS by 2% and 1.9% respectively. Recently published MobileViTv2 architecture removes fusion block and uses linear complexity transformers to perform better than MobileViTv1. We add our proposed fusion block to MobileViTv2 to create MobileViTv3-0.5, 0.75 and 1.0 models. These new models give better accuracy numbers on ImageNet-1k, ADE20K, COCO and PascalVOC2012 datasets as compared to MobileViTv2. MobileViTv3-0.5 and MobileViTv3-0.75 outperforms MobileViTv2-0.5 and MobileViTv2-0.75 by 2.1% and 1.0% respectively on ImageNet-1K dataset. For segmentation task, MobileViTv3-1.0 achieves 2.07% and 1.1% better mIOU compared to MobileViTv2-1.0 on ADE20K dataset and PascalVOC2012 dataset respectively. Our code and the trained models are available at: https://github.com/micronDLA/MobileViTv3

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms

    cs.CV 2025-08 conditional novelty 5.0 of 10

    LOLViT, a GhostNet-based lightweight backbone using adaptive window attention, reports CNN-like CPU speed with MobileViT-level accuracy.

  2. Exploring Superposition and Interference in State-of-the-Art Low-Parameter Vision Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Vision networks whose bottlenecks avoid feature-map interference scale better than MobileNet-style designs at very low parameter counts, and the new NoDepth block demonstrates this on ImageNet.

  3. Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Adding depth-derived geometric priors to a RepVGG backbone improves surgical phase classification on a new nine-phase ESD dataset.

  4. Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications

    cs.RO 2025-07 reject novelty 4.0 of 10

    A synthetic-data-trained pipeline for locating rose flower centers in 2D and estimating their depth in stereo images is described, with in-simulation F1 up to about 96-100%, but real-world detection is below a YOLOv5 ...

  5. TransLPRNet: Lite Vision-Language Network for Single/Dual-line Chinese License Plate Recognition

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A lightweight vision-language transformer with a weakly supervised perspective-correction module reaches about 99% accuracy on modified CCPD license plate benchmarks.

Pith tools