Pith. sign in

REVIEW 2 cited by

Rethinking Mobile Block for Efficient Attention-based Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.01146 v4 pith:LXQHPQVT submitted 2023-01-03 cs.CV

classification cs.CV
keywords attention-basedblockefficientlightweightmodelsmobiledesigneffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper focuses on developing modern, efficient, lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterpart has been recognized by attention-based studies. This work rethinks lightweight infrastructure from efficient IRB and effective components of Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMB) for lightweight model design. Following simple but effective design criterion, we deduce a modern Inverted Residual Mobile Block (iRMB) and build a ResNet-like Efficient MOdel (EMO) with only iRMB for down-stream tasks. Extensive experiments on ImageNet-1K, COCO2017, and ADE20K benchmarks demonstrate the superiority of our EMO over state-of-the-art methods, e.g., EMO-1M/2M/5M achieve 71.5, 75.1, and 78.4 Top-1 that surpass equal-order CNN-/Attention-based models, while trading-off the parameter, efficiency, and accuracy well: running 2.8-4.0x faster than EdgeNeXt on iPhone14.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cascaded Multi-Scale Attention for Enhanced Multi-Scale Feature Extraction and Interaction with Low-Resolution Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CMSA, a grouped multi-head attention with cascaded multi-scale feature fusion, improves accuracy on low-resolution pose estimation and CIFAR classification while using far fewer parameters than prior models.

  2. SCANet: Split Coordinate Attention Network for Building Footprint Extraction

    cs.CV 2025-07 conditional novelty 4.0 of 10

    SCANet with the new Split Coordinate Attention module achieves 91.61% and 75.49% IoU on the WHU and Massachusetts building datasets, slightly beating prior SOTA.

Pith tools