Pith. sign in

REVIEW 8 cited by

EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14756 v6 pith:A7252UFS submitted 2022-05-29 cs.CV

classification cs.CV
keywords efficientvithigh-resolutiondensepredictionattentionmodelsmulti-scaledelivers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

High-resolution dense prediction enables many appealing real-world applications, such as computational photography, autonomous driving, etc. However, the vast computational cost makes deploying state-of-the-art high-resolution dense prediction models on hardware devices difficult. This work presents EfficientViT, a new family of high-resolution vision models with novel multi-scale linear attention. Unlike prior high-resolution dense prediction models that rely on heavy softmax attention, hardware-inefficient large-kernel convolution, or complicated topology structure to obtain good performances, our multi-scale linear attention achieves the global receptive field and multi-scale learning (two desirable features for high-resolution dense prediction) with only lightweight and hardware-efficient operations. As such, EfficientViT delivers remarkable performance gains over previous state-of-the-art models with significant speedup on diverse hardware platforms, including mobile CPU, edge GPU, and cloud GPU. Without performance loss on Cityscapes, our EfficientViT provides up to 13.9$\times$ and 6.2$\times$ GPU latency reduction over SegFormer and SegNeXt, respectively. For super-resolution, EfficientViT delivers up to 6.4x speedup over Restormer while providing 0.11dB gain in PSNR. For Segment Anything, EfficientViT delivers 48.9x higher throughput on A100 GPU while achieving slightly better zero-shot instance segmentation performance on COCO.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SA-Homo: Scale Adaptive Homography Estimation for Scale Variation Scenarios

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    SA-Homo introduces a hierarchical scale-adaptive homography estimation framework with SDBM, MLAC, CSMB, and IHERM modules plus the HMSA dataset that claims robust performance under up to 8x scale discrepancies.

  2. Elastic Attention Cores for Scalable Vision Transformers

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    VECA learns effective visual representations using core-periphery attention where patches interact exclusively via a resolution-invariant set of learned core embeddings, achieving linear O(N) complexity while maintain...

  3. ATT-CR: Adaptive Triangular Transformer for Cloud Removal

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    ATT-CR introduces triangular attention and feature-selected gating to reduce computational cost and cloudy-pixel interference in remote sensing cloud removal.

  4. JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    JetViT uses post-training attention search to hybridize full-attention ViTs with linear and window attention blocks, achieving up to 1.79x throughput gains on high-res images while preserving accuracy on DINOv3 and De...

  5. Enhancing Deep Neural Network Reliability with Refinement and Calibration

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    RefCal uses a supervised contrastive loss to promote refinement alongside calibration in DNN training, reporting better accuracy, refinement, and ECE than baselines on imbalanced CIFAR-100-LT.

  6. ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    ShellfishNet is a new benchmark of 8,691 images across 32 mollusc taxa for evaluating vision models on real-world underwater ecological monitoring tasks including robustness to degradation.

  7. Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms

    cs.CV 2025-08 conditional novelty 5.0 of 10

    LOLViT, a GhostNet-based lightweight backbone using adaptive window attention, reports CNN-like CPU speed with MobileViT-level accuracy.

  8. On The Application of Linear Attention in Multimodal Transformers

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    Linear attention delivers significant computational savings in multimodal transformers and follows the same scaling laws as softmax attention on ViT models trained on LAION-400M with ImageNet-21K zero-shot validation.

Pith tools