Pith. sign in

REVIEW 13 cited by

MobileViTv3: Mobile-Friendly Vision Transformer with Simple and Effective Fusion of Local, Global and Input Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.15159 v2 pith:FJUUVTYT submitted 2022-09-30 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords fusionmodelsblockcreateimagenet-1kade20kbetterdataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

MobileViT (MobileViTv1) combines convolutional neural networks (CNNs) and vision transformers (ViTs) to create light-weight models for mobile vision tasks. Though the main MobileViTv1-block helps to achieve competitive state-of-the-art results, the fusion block inside MobileViTv1-block, creates scaling challenges and has a complex learning task. We propose changes to the fusion block that are simple and effective to create MobileViTv3-block, which addresses the scaling and simplifies the learning task. Our proposed MobileViTv3-block used to create MobileViTv3-XXS, XS and S models outperform MobileViTv1 on ImageNet-1k, ADE20K, COCO and PascalVOC2012 datasets. On ImageNet-1K, MobileViTv3-XXS and MobileViTv3-XS surpasses MobileViTv1-XXS and MobileViTv1-XS by 2% and 1.9% respectively. Recently published MobileViTv2 architecture removes fusion block and uses linear complexity transformers to perform better than MobileViTv1. We add our proposed fusion block to MobileViTv2 to create MobileViTv3-0.5, 0.75 and 1.0 models. These new models give better accuracy numbers on ImageNet-1k, ADE20K, COCO and PascalVOC2012 datasets as compared to MobileViTv2. MobileViTv3-0.5 and MobileViTv3-0.75 outperforms MobileViTv2-0.5 and MobileViTv2-0.75 by 2.1% and 1.0% respectively on ImageNet-1K dataset. For segmentation task, MobileViTv3-1.0 achieves 2.07% and 1.1% better mIOU compared to MobileViTv2-1.0 on ADE20K dataset and PascalVOC2012 dataset respectively. Our code and the trained models are available at: https://github.com/micronDLA/MobileViTv3

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EMOv2: Pushing 5M Vision Model Frontier

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 5M-parameter backbone with shared-weight spanning window attention sets new accuracy records across classification, detection, and generation benchmarks.

  2. Improving Accuracy and Generalization for Efficient Visual Tracking

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SiamABC, a Siamese tracker with a dual-search-region, a fast filtration layer, and backward-free test-time adaptation, improves out-of-distribution tracking while running at 100 FPS on a CPU.

  3. MobileMamba: Lightweight Multi-Receptive Visual Mamba Network

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MobileMamba achieves 73.6 to 83.6% ImageNet top-1 accuracy across six model sizes, with GPU throughput up to roughly 21 times that of LocalVim, by combining wavelet-enhanced Mamba and multi-kernel depthwise convolutio...

  4. Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms

    cs.CV 2025-08 conditional novelty 5.0 of 10

    LOLViT, a GhostNet-based lightweight backbone using adaptive window attention, reports CNN-like CPU speed with MobileViT-level accuracy.

  5. Exploring Superposition and Interference in State-of-the-Art Low-Parameter Vision Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Vision networks whose bottlenecks avoid feature-map interference scale better than MobileNet-style designs at very low parameter counts, and the new NoDepth block demonstrates this on ImageNet.

  6. Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Adding depth-derived geometric priors to a RepVGG backbone improves surgical phase classification on a new nine-phase ESD dataset.

  7. Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A spike-driven Transformer trained with integer activations reaches 86.2% top-1 on ImageNet, the highest reported accuracy for a directly trained spiking network at this scale.

  8. Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications

    cs.RO 2025-07 reject novelty 4.0 of 10

    A synthetic-data-trained pipeline for locating rose flower centers in 2D and estimating their depth in stereo images is described, with in-simulation F1 up to about 96-100%, but real-world detection is below a YOLOv5 ...

  9. TransLPRNet: Lite Vision-Language Network for Single/Dual-line Chinese License Plate Recognition

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A lightweight vision-language transformer with a weakly supervised perspective-correction module reaches about 99% accuracy on modified CCPD license plate benchmarks.

  10. A Multimodal In Vitro Diagnostic Method for Parkinson's Disease Combining Facial Expressions and Behavioral Gait Data

    cs.CV 2025-06 reject novelty 4.0 of 10

    A multimodal deep-learning pipeline fusing gait and facial-expression features reports perfect Parkinson's diagnosis accuracy on a small, unreleased dataset.

  11. RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations

    cs.CV 2024-12 conditional novelty 4.0 of 10

    RecConv recursively decomposes feature maps into multiple scales with shared small-kernel depthwise convolutions to grow the effective receptive field to k times 2^ell at roughly constant FLOPs, yielding the RecNeXt b...

  12. LUIEO: A Lightweight Model for Integrating Underwater Image Enhancement and Object Detection

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A lightweight multi-task network performs underwater image enhancement and object detection jointly, using a physical scattering model for self-supervision, and reports higher mAP50 than YOLOv8 on RUOD.

  13. EM-Net: Gaze Estimation with Expectation Maximization Algorithm

    cs.CV 2024-12 reject novelty 3.0 of 10

    EM-Net, a MobileNetV3-based gaze estimator with a Swin-style attention branch and an EM refinement module, reports angular-error gains of 0.08 to 0.24 degrees over GazeNAS-ETH using 50% training data.

Pith tools